mirror of
https://github.com/heygen-com/hyperframes.git
synced 2026-09-07 18:26:17 +00:00
eba8a0fa2913cfac6ccffdf04ed28558a9d4f9b2
322
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0045e8b3c5 |
fix(producer): stop router from mutating process.env for cross-render state
HF_DE_PARALLEL_STREAM was restored on every exit path, but the producer server allows concurrent renders in one process — a router-eligible job's mutation was still visible to an unrelated render already executing during that window. Thread the router's decision as a per-render local instead of a global env var; HF_DE_PARALLEL_STREAM stays as the manual opt-in. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
a23e63326d |
fix(producer): restore HF_DE_PARALLEL_STREAM on every render exit path
Address unanimous review feedback on PR #2095 (Miga, Magi REQUEST_CHANGES, Rames D Jusso): the router's process.env.HF_DE_PARALLEL_STREAM mutation was only cleared on the DE self-verify-retry branch — every other exit (happy path, any non-DE-verify error, abort) left it set, leaking the parallel-streaming opt-in into the next render sharing the same process (a regression/benchmark harness, or any batch host). Capture the prior value before mutating and restore it in the outer `finally` (executeRenderJob's own top-level try/finally, which runs on every exit path by construction), not just the narrow retry branch. Note `deParallelRouter` and the new `deParallelStreamEnvBefore` had to move above the outer `try` — a `let` declared inside `try {}` is not visible in the sibling `finally {}` block in JS, so the original placement alongside the other DE state would not have compiled once referenced from the finally. Also: renamed the shared `preInversionWorkerCount` local to `preRoutingWorkerCount` (Miga + Rames nit — it now serves both the inversion and the router), and pinned worker count 3 explicitly overrides calibration by design (documented per Miga/Rames's question, not a bug). Verified end-to-end (not just unit tests): ran two executeRenderJob calls back-to-back in one process, router-eligible then not — env is restored to undefined after render 1 and stays clean through render 2, the exact leak scenario the reviews described. 117 orchestrator tests pass (3 new, covering the restoreEnv primitive directly). |
||
|
|
0393ba5be2 |
feat(producer): default-off router for verified parallel drawElement
Promotes the opt-in HF_DE_PARALLEL_STREAM mechanism (#2056) into the auto-routing decision, gated behind its own default-off flag (HF_DE_PARALLEL_ROUTER). This is the next step from the 2026-07-08 parallel-DE benchmark verdict: par3/single 1.16-1.36x on real-work comps >=2,000 frames, no comp anywhere losing to single-worker. shouldPreferParallelDrawElement mirrors shouldPreferSingleWorkerDrawElement (#2026) but takes priority over it when both are eligible — its higher default threshold (HF_DE_PARALLEL_MIN_FRAMES=2000 vs the inversion's 900) means it only ever picks up the long tail the inversion's own benchmark didn't cover. Fixed at 3 workers (benchmark-validated; not calibration- derived), same shape as the inversion pinning to a fixed 1. resolveParallelRouterRetryPlan mirrors resolveInversionRetryPlan for the self-verify-failure rollback path: falls back to the ordinary (non-DE) parallel-disk path at the pre-router worker count. The caller must clear HF_DE_PARALLEL_STREAM before recomputing useStreamingEncode or the retry would keep resolving to the parallel-streaming shape. New telemetry (de_parallel_router, de_pre_router_workers) tags which render used the router, separate from de_worker_inversion — needed before the planned telemetry soak can segment revert-rate and de_verify_min_db to the parallel cohort specifically; today there's no way to tell those apart from ordinary single-worker DE renders. Verified end-to-end: HF_DE_PARALLEL_ROUTER=true on a 2,381-frame comp resolves to 3 workers with 3 separate drawElement sessions and renders successfully; without the flag, behavior is unchanged (falls through to the existing single-worker inversion, workerCount=1) — no regression to current production routing. 114 orchestrator tests pass (15 new). |
||
|
|
381887541f |
fix(engine,producer): fix quadratic dedup rescan, correct race justification
Address two max-effort code-review findings on PR #2056 not covered by the earlier review-gap commit: - captureFrameToBufferPipelined's static-dedup reuse branch never advanced session.lastEncodeResultFrame, unlike its sibling real-capture branches. The gap-check window is computed from that watermark, so every consecutive reuse in a static run rescanned an ever-widening window instead of just the newest frame — O(n^2) total work over a long static stretch instead of O(n). - The "single-threaded, no race" justification on the shared parallelGuard closure was wrong: the guard has real internal await points (recapture, PSNR) between reading and writing its sizes/absFloor/acceptedSmall state, so concurrent workers' calls do interleave there (confirmed). Replaced with the actual reason it's safe: absFloor only ratchets down, sizes is append-only and order-independent for the median, and acceptedSmall's fast path re-validates by exact byte-equality regardless of which worker wrote the reference buffer. |
||
|
|
2dbe958a49 |
fix(engine,producer): close review gaps in parallel drawElement streaming
Address PR #2056 review feedback: - Fix totalFrames progress inflation for interleaved tasks — divide each task's span by its frameStride to match the actual per-worker frame count (captureFrameRange steps by stride), instead of summing raw endFrame-startFrame which double(N)-counts interleaved tasks. - Attach a no-op .catch to each frame's pipelined encodeResult at kick time so an abandoned promise (loop exits early on abort/error before draining it) can't surface as an unhandled rejection during teardown. - Document why the pipelined branch's stride=1 path is validation-only in production (HF_DE_PARALLEL_STREAM always uses interleaved distribution) so a future refactor doesn't unknowingly widen it. - Comment the intentional single shared parallelGuard/parallelStats across workers (safe single-threaded, better rolling-median signal). |
||
|
|
81a4f04360 | docs(producer): disambiguate compositor frame scheduling from the beginframe capture mode | ||
|
|
b3493a7b61 |
feat(engine,producer): verified interleaved parallel drawElement streaming (opt-in)
Step 2 of the DE engagement plan: multi-worker drawElement capture through the streaming encoder, with the full runtime self-verification net riding along — the confinement rule that kept the parallel clamp in place is now satisfied on this path. Opt-in via HF_DE_PARALLEL_STREAM=true; default routing (including the #2026 single-worker inversion) is unchanged. Mechanism: - distributeFramesInterleaved + WorkerTask.frameStride: worker i captures frames i, i+N, i+2N... — seek-based capture makes stride free and the ordered writer's reorder window shrinks from totalFrames/N to N (contiguous chunks serialize workers behind the writer). - Depth-2 pipelined worker-encode produce in the parallel worker loop (the same shape as the sequential loop; frame k's in-page encode overlaps k+stride's produce). HF_DE_PAR_DEBUG=1 traces the first frames per worker. - Drain guard extracted to createDrainFrameGuard (session-parameterized): every parallel frame gets the SAME blank-guard + PSNR self-verify as the sequential drain, against its owning worker's pre-injection ground truth (all sessions arm identical sample indices from CaptureOptions.compositionDurationSeconds). - FrameReorderBuffer.abort(err): a failed worker (e.g. verification error) rejects all parked and future waiters — without this, peers park forever in waitForFrame and the pool (which awaits ALL workers before surfacing errors) deadlocks. Found by the verify-trip test; unit-tested. - The typed DrawElementVerificationError is preserved past the pool's error-string flattening so the orchestrator's verify-retry recognizes it. - Static-dedup stride hazard fixed: lastEncodeResult reuse now requires EVERY frame in (lastEncodeResultFrame, i] to be predicted-static (sequential capture reduces to the old has(i) check). - Workers get separate browser PROCESSES under the flag: pages co-tenant in one browser starve non-active pages of BeginFrames on the paint-wait path (measured 86s vs 30s on a 3,245-frame rAF comp). Validation: - Happy path W3: verify samples pass across workers (4x inf on the 2,381f comp), output vs single-worker DE = 59.3dB (encode noise floor) — the interleave + dedup-stride produce identical pixels. - Verify-trip (marginal comp + HF_DE_VERIFY_MIN_DB=45): fails at frame 649 (32.2dB < 45), peers abort instead of deadlocking, whole render retries via parallel screenshot, RENDER_OK in 42.6s. - Canary suite 7/7 with the flag off (default paths untouched); producer orchestrator tests 99/99; engine suite 909 passed (14 pre-existing main failures, stash-A/B verified); reorder-buffer abort unit tests. Perf note: capture-only parallel speedup measured 1.38x (W2) / 1.52x (W3) over single-worker DE in the spike; end-to-end numbers on this machine are currently noisy (separate-browser init overhead + bench load) — clean benchmarks before any default routing change. The flag stays explicit opt-in; promoting it into the router replaces the #2026 W=1 pin for the same cohort. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
8ba3c33915 |
fix(engine,producer,cli): close review gaps in sub-timeline fail-fast
Address PR #2045 review feedback: - Share a SubTimelineWaitOutcome type (engine) end-to-end instead of widening to string across CapturePerfSummary / RenderPerfSummary / telemetry, so the three layers can't drift. - Dedupe scriptLoadFailures on push — a 4xx response and its trailing requestfailed both recorded the same URL, doubling the failed-URL list in the fail-fast warning. - Thread the sub-timeline-wait outcome into render_error (not just render_complete): a render that fail-fasts and then fails downstream (pollVideosReady, extract, encode) previously dropped this signal on the floor. dedupPerfs is now function-scoped so the catch path can read it, same treatment as the existing captureAttempts array. |
||
|
|
54359f3d6a |
fix(engine): fail-fast the sub-composition timeline wait when a script 404s
pollSubCompositionTimelines waits for every [data-composition-id] host to
register window.__timelines[id]. When the script carrying that registration
fails to load (404 / request failure), the registration can never arrive —
but the poll still burned the full playerReadyTimeout (45s), then warned and
shipped a silently animation-less render. Wild scale: the capture-setup
histogram over 30 days of local renders decays smoothly (402/503/364/282/191
per 5s bucket) then spikes to 705 at the 45s bucket — ~1,000 renders/month
across 402 distinct users, ~15 user-hours of pure waiting.
- Sessions now record failed SCRIPT resources (requestfailed + HTTP>=400
response, listeners that already existed for diagnostics) in
session.scriptLoadFailures.
- pollSubCompositionTimelines takes a failure getter and cuts the wait to a
2s grace once any script failed, with a loud warning naming the URL(s).
Late-registering fetch-async comps are unaffected: no script failure means
the full timeout still applies, and a registration landing inside the
grace window still wins (tested).
- Outcome telemetry: session.subTimelineWaitOutcome ("ready" | "timeout" |
"script_failure") -> CapturePerfSummary -> RenderPerfSummary.subTimelineWait
(worst across sessions) -> render_complete sub_timeline_wait, so the wild
rate becomes directly trackable instead of setup-histogram forensics.
Validation: the discovery comp (0768f038, its animations.js unreachable)
drops from ~72s to 23.1s total — poll cut at 2.1s with the script named;
healthy comp reports "ready". Canary suite 7/7 (PSNRs identical). 4 new
poll unit tests; engine suite 907 passed (14 failures are PRE-EXISTING on
main at v0.7.42 — 18 fail on a clean checkout, stash A/B verified).
tsc/oxlint/oxfmt clean.
Corpus note: 258/1,762 corpus comps (14%) reference local scripts missing
from the corpus fetch — their historical eval INIT timings measured this
timeout, not the engine. Capture-stage ratios remain valid (both paths paid
it equally).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
5ebc5bb10f |
fix(producer): scope per-instance variables for repeated sub-composition mounts (#2070)
#2066 fixed sub-composition data-variable-values on the render path for a single mount, but the reusable-template pattern from #2064 (the same sub-comp mounted multiple times with different values) still diverged from preview/snapshot: every mount shared one __hfVariablesByComp key and one CSS scope selector, so the last mount's values clobbered the earlier ones and all-but-one instance rendered blank. The producer now assigns per-instance runtime composition ids (assignBundledRuntimeCompositionIds) and threads hostIdentityMap into the shared inliner, mirroring the preview bundler. The shared inliner's default buildScopeSelector already scopes by the runtime id, and timelines remap to it via the scoping proxy, so each instance's variables, CSS, and timeline land under its own id. Pixel-verified end to end: two mounts of one sub-comp with different data-variable-values now render their own content (green CARD_A / blue CARD_B), matching snapshot; single-instance behavior is unchanged. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
41ad5b4690 |
fix(engine): inject sub-composition variables on the render path (#2066)
render left window.__hyperframes.getVariables() empty inside every
sub-composition mounted via data-composition-src, so each instance rendered
its declared JS defaults instead of the per-instance data-variable-values.
preview/snapshot injected them correctly, so the composition looked right in
every authoring/QA surface and then rendered wrong content silently (exit 0).
Any template-library workflow (reusable sub-comp scenes parametrized per
video) shipped placeholder/default text in the final MP4.
The plumbing already existed on main: htmlCompiler passes
readVariableDefaults/parseHostVariables and populates result.variablesByComp,
and the CSS-custom-property path (emitRootCompositionVariableStyles) reaches
the render. But the render compiler emitted only the CSS vars and never the
JS table window.__hfVariablesByComp that the scoped getVariables reads, while
the preview bundler (htmlBundler) did -- so getVariables() returned {} only
during render.
Fix, so the paths cannot drift again: buildVariablesByCompScript, colocated
with the reader in compositionScoping.ts and shared by both compile paths.
htmlBundler now calls it instead of an inline string; htmlCompiler injects it
before the inlined sub-comp scripts, using the already-populated
result.variablesByComp.
Verified end-to-end: a sub-comp painting its background from a color variable
now renders the injected value under render, matching snapshot; previously it
rendered the default. 3 new producer tests; 89 htmlCompiler + core-compiler
tests pass.
Closes #2064.
|
||
|
|
e2c88ef689 |
fix(core,producer): composition CSS variables reach the render path at eval time
Live testing of the compile-time variable emission surfaced four gaps: - The producer render path never emitted the compile-time stylesheet (only the preview bundler did), so eval-time reads — GSAP .from immediateRender, top-level getComputedStyle — saw undefined vars in rendered output. The producer's inlineSubCompositions now calls the shared emitRootCompositionVariableStyles and passes the variable hooks. - --variables overrides weren't visible at eval time. They now thread from the orchestrator / distributed plan through compileStage into the emitted rules (window.__hfVariables still covers script reads). - Per-declarer rules anchored on data-composition-id, which two inlined instances of one sub-composition share — instance A's rule restyled instance B, and a rule directly on the declarer defeated the host's inherited data-variable-values. Rules now anchor on per-instance data-hf-var-scope markers and layer nearest-host values over declared defaults, mirroring the runtime loader. - Emission ignored authored CSS; a declared default now yields to a var already defined in an authored <style> block (define-if-absent, matching the runtime injection). Also: the figma importer emits background-color (longhand) for solid fills. GSAP backgroundColor tweens cannot read a var() through the background shorthand — its pending-substitution longhands serialize empty, so .from captured nothing and settled on transparent (pre-existing GSAP interaction, reproduced with no composition variables involved). Validated live: eval-time default + override, .from + override, two-instance host branding, authored :root precedence, SDS brand-loop pixel parity. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5b9b71df25 |
fix(producer): suppress GSAP call side effects during render seeks (#2037)
* fix(producer): suppress GSAP call side effects during render seeks * fix(core): preserve GSAP root render nudge safely |
||
|
|
924727a0b4 |
feat(producer,cli): drawElement priority inversion — single-worker streaming over auto-parallel (#2026)
* feat(producer,cli): drawElement priority inversion — single-worker streaming over auto-parallel clamp:parallel eats 50% of local renders (1,326/fortnight; DE engagement stuck at 3.8%) by routing multi-worker renders to unverified screenshot capture. Benchmarks (2026-07-08, 4 comps x W1/W2/W3/W5) show that above the ~900-frame amortization crossover, single-worker VERIFIED drawElement streaming beats screenshot-parallel at EVERY worker count (2,380f: 66s vs 109-127s; 3,600f: 33s vs 39-56s; parallel scaling flattens past W2), while below it DE's fixed init cost loses by <=2.2s. - shouldPreferSingleWorkerDrawElement (exported predicate + 7 unit tests): inverts an AUTO-resolved multi-worker render to workerCount=1 when the comp matches the benchmarked configuration — default-on DE (darwin hardware clamp upstream), no compile gate, no forced-screenshot hint, mp4 output, single-worker streaming eligible, and totalFrames >= HF_DE_SINGLE_MIN_FRAMES (default 900; 0 disables). Explicit --workers N is always honored. - Inverted renders keep the probe session and land on the worker-encode streaming drain — the ONLY path with runtime self-verification, so this moves ~40% of previously-clamped renders onto the verified fast path. Comps that later hit an init-time gate (~1.5% of local renders) render single-worker screenshot streaming; accepted trade. - Telemetry: de_worker_inversion on render_complete (orchestrator -> perfSummary.workerInversion -> CLI), plus the worker_resolution observability checkpoint now records deWorkerInversion. Validation: e2e matrix on 2,381f comp — auto->5 workers inverted to 1, DE verified 4x inf PSNR, RENDER_OK; short comp (360f) auto stays 5-worker; explicit WORKERS=3 honored; HF_DE_SINGLE_MIN_FRAMES=0 disables. Canary suite 7/7 (PSNRs identical). renderOrchestrator tests 86/86. tsc/oxlint/oxfmt clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(producer): review fixes — inversion routing guards, calibration skip, retry revert Max code-review round on the inversion (13 confirmed findings): - Streaming spawn-failure disk fallback now clamps default-on drawElement (deClampReason=disk_path, DE-mode probe closed) exactly like the pre-capture clamp — previously it carried useDrawElement=true onto the unverified disk path, the hole the verified-path confinement exists to close, newly reachable for every inverted render. - Predicate gained the routing knowledge it was blind to: layered/HDR and shader-transition comps (drawElement never runs there), supersampling (deviceScaleFactor>1 init gate), a probe session whose init gates already disengaged DE, and the PRODUCER_EXPERIMENTAL_FAST_CAPTURE=true explicit parallel-DE opt-in (honored like --workers N). - Eligibility is evaluated BEFORE capture calibration and skips it when the inversion pins workers to 1 regardless of the estimate — the throwaway calibration browser + sample captures cost ~41s on the 2,381-frame benchmark comp (auto render: 111.6s -> 70.1s total). - Self-verify retry reverts the inversion: the re-render returns to the pre-inversion parallel screenshot path (disk) instead of single-worker screenshot streaming, the slowest shape for exactly the comps drawElement damages. - HF_DE_SINGLE_MIN_FRAMES="" (set-but-empty) now falls back to the 900 default instead of aliasing the 0 kill switch. - Timeout advisory uses the RESOLVED worker count — an inverted render that times out no longer prints "Retry with --workers 1" (the configuration that just failed). - Telemetry: deWorkerInversion recorded in capture observability (failed renders are attributable), emitted as literal false when not fired (queryable denominator), and the drawElement perf input shape is one exported DrawElementPerfInput type instead of three copies. - Tests: requestedWorkers undefined (the value production actually passes) + the four new predicate guards; 91/91. Validation: e2e auto render — calibration skipped (deInversionEligible), inversion fires, DE verified 4x inf, total 70.1s (was 111.6s); HF_DE_SINGLE_MIN_FRAMES=0 restores calibration + parallel; canary suite 7/7 (PSNRs identical); tsc/oxlint/oxfmt clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(producer,cli): review round 2 — loss-cohort telemetry, retry-plan helper, boundary tests - de_worker_inversion is now a tri-state string ("inverted" | "reverted" | "none") instead of a boolean: the self-verify retry marks the render "reverted" rather than resetting to false, so the dashboard can segment the lost-inversion cohort first-class instead of inferring it from deSelfVerifyFallback + frame-count joins (james-russo #1). - The retry rollback is extracted to resolveInversionRetryPlan (pure, exported) with unit coverage: pre-inversion worker-count restore, streaming re-resolution (multi-worker retry -> disk), "reverted" state, null when never inverted (james-russo #2). - WOULD_RESOLVE_MULTI_WORKER named constant replaces the bare sentinel 2 (james-russo #5); minFrames: -1 boundary case added (miga #3). 94/94 renderOrchestrator tests; tsc/oxlint/oxfmt clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(producer,cli): emit de_pre_inversion_workers for the parallel counterfactual The ramp-down decision needs "did DE beat the parallel render it displaced", not just "did DE beat single-worker screenshot". Emitting the worker count the auto-resolution chose BEFORE the inversion pinned it to 1 makes the parallel counterfactual computable per render (screenshot ms/frame from the verify samples / W x the measured parallel-efficiency curve). Set only when the inversion fired. Smoke: 2,381f auto render -> de_worker_inversion="inverted", de_pre_inversion_workers=5, mode=drawelement, verify armed 4. 99/99 tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
00b96d2eaa |
fix(producer): fall back to copying extracted frames when symlink hits EPERM
* fix(producer): fall back to copying extracted frames when symlink hits EPERM materializeExtractedFramesForCompiledDir stages each video's extracted frames into the compiled dir via a single symlink (the in-process renderer's default; distributed plan() already copies via materializeSymlinks). Windows without Developer Mode (or Administrator) cannot create symlinks and rejects with EPERM, so high/standard-quality renders failed there — while draft quality worked because it avoids the symlinked-cache path entirely. Fix: a new stageExtractedFrameDir helper catches EPERM/EACCES from symlinkSync and falls back to the same recursive cpSync the materializeSymlinks path already uses. The extra disk is far better than a hard render failure on a default Windows configuration. Non-permission errors (ENOSPC, etc.) still propagate so real failures aren't masked as silent copies. Extracting the helper also keeps the main function under the complexity gate. Test: two new cases via the injected fileSystem — symlinkSync throwing EPERM triggers exactly one recursive cpSync (frames still remapped under compiledDir), and a non-permission error (ENOSPC) rethrows without falling back to copy. Full renderOrchestrator suite (81) passes. * fix(producer): recover from a stale dangling frame-symlink (EEXIST) Follow-up to this PR's EPERM copy fallback, from a further Windows report: the symlink fails with EEXIST after the extraction cache is GC'd. A prior render's symlink at the compiled linkPath dangles once its target is removed; the caller's existsSync() guard follows the dead link and reads it as absent, so staging runs again, but the link file still exists and symlinkSync collides with EEXIST -> the render hard-fails. Catch EEXIST in stageExtractedFrameDir, clear the stale entry (rmSync), and re-stage (link, or copy on EPERM/EACCES). Factored the link-or-copy into a helper reused by both the first attempt and the retry. rmSync is an optional injected fs method (default fs supplies it; only the EEXIST path calls it). New unit test covers the dangling-symlink recovery. * fix(producer): widen symlink fallback to UNKNOWN and cover EEXIST on copy path Addresses review nits on the frame-staging fallback: - Widen the symlink no-privilege catch from EPERM/EACCES to also include UNKNOWN (some Windows builds surface a symlink privilege denial as UNKNOWN). - Wrap the EEXIST stale-entry recovery around BOTH staging branches, not just the symlink one: after #2025 Windows uses the eager cpSync path, which can collide with a dangling symlink left by a prior Linux run — now it clears the stale entry and re-stages either way. - Emit a one-time INFO log when symlinking degrades to copying, so a heavier Windows render is self-explanatory. |
||
|
|
1aa39d4653 |
fix(producer): copy extracted frames on Windows to avoid symlink EPERM
The local render path (renderOrchestrator → runExtractVideosStage) materialized each video's extracted frames into the compiled dir via symlinkSync, with no materializeSymlinks flag. On Windows without Developer Mode/Administrator, symlinkSync throws EPERM, so local video renders failed at the video_extract stage (users worked around it with materializeSymlinks patches / snapshot-frame hacks). The distributed plan() path already copies (materializeSymlinks: true). Pass materializeSymlinks: shouldCopyExtractedFrames(process.platform) at the local caller — copy on win32 (symlinks unavailable), symlink elsewhere (cheap). New pure shouldCopyExtractedFrames() helper + unit tests. |
||
|
|
337d0b51bc |
refactor(engine,producer): adopt requestPaint contract, retire autoAlpha rewrite (#2021)
* refactor(engine,producer): adopt requestPaint contract, retire autoAlpha rewrite crbug 529829538 was closed "working as intended": the html-in-canvas API's contract is mutate -> canvas.requestPaint() -> await the canvas paint event -> drawElementImage, which refreshes the subtree's paint records including compositor-applied properties. Verified on the pinned 151 floor and 152 canary: root opacity, root filter, nested group opacity, and child transforms (incl. will-change-promoted) all capture exactly; the root element's own TRANSFORM is the one property still never baked. - Paint invalidation: all three paint-wait sites (serial capture, worker produce, batch produce) now call canvas.requestPaint() when available and fall back to the __hf_de_tick sentinel background toggle on builds without it. The 250ms unsynchronized-draw safety net is unchanged. - Root-opacity ratio correction REMOVED (all three draw sites + base-opacity recording at injection). Since 151 the paint wait bakes current root opacity into the snapshot as pixel alpha, so the ratio correction DOUBLE-APPLIED animated root fades: a root-fade A/B tripped the runtime self-verify at 30.1dB (frame 24, ~0.92 expected vs ~0.85 rendered). Post-removal the same comp self-verifies at inf and matches the screenshot render at PSNR=inf. The root TRANSFORM correction stays — verified still required. - autoAlpha rewrite machinery DELETED: the opt-in opacity->autoAlpha tween rewrite (default-off since the retraction fix; measured ~28dB damage on comps whose fades it touched), its flush-time transparent-target hiding, the __HF_FAST_CAPTURE_AUTOALPHA__ flag plumbing, and the deferral-time retract/re-assert dance. The stub keeps tween-target tracking (3D projection + at-risk scans depend on it). Validation: canary suite 7/7 with PSNRs identical to baseline (58.30 / 43.13 / 54.15 dB); root-fade A/B PSNR=inf vs screenshot; engine suite 905 passed (1 pre-existing color-grading failure); tsc/oxlint/oxfmt clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(engine,producer): review fixes — gate opacity correction by paint mechanism Max code-review findings on the requestPaint adoption: - Root-opacity ratio correction RESTORED, gated per frame on how the paint was produced: it applies on BeginFrame (sync=false) captures and on builds without canvas.requestPaint() — the two paths where the snapshot holds the root's load-time opacity — and is skipped only on requestPaint-driven paints, where the snapshot bakes the current opacity and the ratio double-applies (the proven 30.1dB root-fade failure). Base opacity is recorded at injection again. - Invalidation extracted to a page-scope helper (__hfDeInvalidate, installed by injectDrawElementCanvas) shared by all three paint-wait sites: sentinel toggle ALWAYS (a paint is guaranteed even if requestPaint elides one on a clean subtree) + requestPaint() in a try/catch (a throwing implementation degrades to sentinel-only instead of rejecting the capture). Returns whether requestPaint ran, feeding the opacity-correction gate. Also removes the triplicated inline block and its three anonymous `as T` casts. - HF_FAST_CAPTURE_AUTOALPHA now logs a retirement warning instead of being a silent no-op (the deleted rewrite's comment documented it as an operator escape hatch). - Batch producer docstring updated (still described the tick-toggle-only paint wait); stub tween observer reshaped to a void fn (observeTweenCall) so no arg-rewriting seam survives. Validation: canary suite 7/7 (58.30/43.13/54.15dB, d95f20b6 clean); root-fade A/B self-verify 4x inf + whole-video PSNR=inf; engine suite 905 passed (1 pre-existing); tsc/oxlint/oxfmt clean; stub regenerated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: WaterrrForever <miao.yang@heygen.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b26c27576b |
feat(engine,producer,cli): verify video comps via deferred DE init + capture p50 (#2015)
* feat(engine,producer,cli): verify video comps via deferred DE init + capture p50 Closes the two biggest gaps in the first day of v0.7.38 wild data: 88% of drawElement renders (video comps initialized via probe sessions) ran with self-verification unarmed, and speedup was measurable on only 3 of 76 renders. - Deferred drawElement init: probe sessions initialize before video extraction, so they have no frame injector — ground-truth screenshots would capture black <video> boxes, and verification skipped the whole comp. DE init now stops after the gates for injector-less video comps (deInitDeferred; autoAlpha flag retracted in case no path completes it) and completeDeferredDrawElementInit finishes verification + canvas injection + worker-encode at capture time, once prepareCaptureSessionForReuse has attached the injector. Validated end-to-end: a probe-path video comp now arms 4 ground-truth frames with real video pixels (3x inf + 64.7dB) and renders drawElement verified. - capture_p50_ms: per-frame capture durations are sampled (capturePerf.frameMs; batch frames get the batch mean) and the median ships as CapturePerfSummary.p50TotalMs -> RenderPerfSummary.captureP50Ms -> render_complete capture_p50_ms. Unlike capture_avg_ms it is immune to first-frame warmup and stage-setup amortization — smoke: avg 15ms vs p50 8ms on the same render, p50 matching the measured steady-state floor. Dashboard speedup tiles can drop their frame-count floor once this ships. - video_count on render_complete: segments speedup by video-injection comps (whose per-frame gain is legitimately lower) vs pure-graphics. Canary suite 7/7; engine suite 905 passed (1 pre-existing upstream failure); tsc/oxlint/oxfmt clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(producer): complete deferred drawElement init on the disk capture path Review (miga): a probe-initialized video comp falling back to the disk path kept deInitDeferred and silently stayed in screenshot mode — a regression for PRODUCER_EXPERIMENTAL_FAST_CAPTURE=true renders that previously ran drawElement there. Complete the deferred init on the sequential disk path under the same explicit-opt-in test the orchestrator clamp uses; default-on renders stay on the screenshot baseline (this path has no drain-time self-verification, per the #1998 confinement rule). Validated: video comp + PRODUCER_ENABLE_STREAMING_ENCODE=false + explicit opt-in logs "(deferred drawElement init)" completion on capture_disk and renders correct video pixels. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
a2a80d5a5c | fix(producer): mix muted browser media as silent for preview-render parity (#1969) | ||
|
|
1005703441 |
feat(engine,producer,cli): drawElement release telemetry on render_complete (#2002)
Default-on drawElement ships with a runtime self-verification net (#1998); this makes its in-the-wild behavior observable. Every render_complete event now answers: which capture mode actually ran, why drawElement disengaged when it did (compile gate / producer clamp / engine init gate), whether the self-verify net fired and why, and how much margin verification had. Follows the static-dedup telemetry pattern: engine session fields → CapturePerfSummary → RenderPerfSummary.drawElement → snake_case props on render_complete. New event props: de_capture_mode, de_compile_gate, de_clamp_reason, de_gate_reason, de_worker_encode, de_verify_armed, de_verify_checked, de_verify_min_db (margin above the 32dB threshold — drift here is the early-warning signal before fallbacks spike), de_verify_init_ms, de_self_verify_fallback, de_fallback_reason, de_blank_suspects, de_blank_deterministic_accepts, de_blank_recaptures, de_boundary_frames, de_ncpr_fallbacks. Validated end-to-end on live renders: drawelement path reports mode/verify counters/minDb/init cost; a blur-gated comp reports mode=screenshot + gate_reason=css_effect:filter; a forced verification failure reports self_verify_fallback=true + fallback_reason=psnr. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
906c8d04f8 |
fix(cli): localize remote assets before validate so it matches render (#2001)
validate served the composition over a loopback origin and let headless
Chrome fetch remote <img crossorigin>/@font-face assets cross-origin, while
the render pipeline downloads them to disk first. Buckets whose CORS
allowlist omits the loopback origin then failed the CORS-mode request with a
false net::ERR_FAILED that never occurs in the real render, pushing authors
(and agent pipelines) to delete crossorigin — which disables WebGL
color-grading/shaders for that asset.
Reuse producer's localizeRemote{Media,Image,FontFace}Sources in validate,
downloading into a temp dir served as an extra static-server asset root
(project dir untouched, cleaned up after). validate now matches render.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
ec06f4bf89 |
feat(engine,producer): drawElement fast-capture default-on with runtime self-verification safety net (#1998)
* feat(engine,producer): drawElement fast-capture default-on with runtime self-verification safety net Flip useDrawElement + worker-encode defaults on (HF_DE_BATCH default 4), clamped in resolveConfig to hosts where drawElement can engage (macOS + hardware-GPU browser) so page-side shader compositing is untouched everywhere else; explicit env opt-in keeps attempt-and-gate semantics. Safety net makes default-on safe: the compile/init gates catch predictable incompatibility; this catches the intermittent residue no static analysis can see (stale paints, dropped background images, transient blank frames). - engine: captureDeVerificationFrames — K=4 (HF_DE_VERIFY) ground-truth screenshots at init, after gates + armStaticDedup, BEFORE canvas injection (post-injection screenshots show the canvas bitmap, not the DOM). Runs the video-injection hook per sample; double-captures so rAF-driven text counters settle (a single immediate screenshot captures stale text and false-positives). Skips png, <10 frames, implausible __hf.duration (infinite-repeat GSAP sentinel). - producer: guardFrame on both worker-encode drains — rolling-median blank guard with retry-once at drain (byte-identical retry ⇒ deterministic dark frame, accepted; retry save/restores the static-dedup anchor) + ffmpeg PSNR self-verify vs ground truth (HF_DE_VERIFY_MIN_DB, default 32dB; natural agreement ≥45dB, damage ≤25dB). Breach dumps the frame pair to tmpdir and throws DrawElementVerificationError. - orchestrator: one-shot retry — on verification error the whole render re-runs with forceScreenshot (slower, never wrong); telemetry flag deSelfVerifyFallback. - tooling: de-canary-suite.sh (7-comp release gate with expected verdicts), de-gatecheck.sh (init-only corpus routing classifier), we-render.mjs. Validated: canary suite 7/7; 611-comp routing sample 54% drawelement / 37.5% gated / 8.3% comp-defect; 12/12 risk-band renders clean on bare defaults (48/48 verify samples); engine suite 888 passed; caught two real intermittent damage classes in the wild (background-image drop, root-props offset) that previously shipped silently. Kill switches: PRODUCER_EXPERIMENTAL_FAST_CAPTURE=false, HF_DE_WORKER_ENCODE=false, HF_DE_BATCH=0, HF_DE_VERIFY=0. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(engine,producer): harden the drawElement self-verification net (max code-review findings) 15 confirmed findings from the adversarial review of the default-on flip; the load-bearing five: - Ground-truth capture no longer scrubs GSAP state: seek(0) + forced frame FIRST (lazy .from()/overlap tweens record start values on first seek — mid-timeline scrubs corrupted them for the whole render, and since DE frames and truth shared the corruption, PSNR passed on damaged output), then ascending even-spread fractions, page left at frame 0. - Default-on drawElement is confined to the verified path: resolveConfig requires worker-encode (the drain that runs the net), the orchestrator disengages the default when the render takes the disk path or parallel capture (no drain verification there), and closes a drawElement-initialized probe session rather than letting the unverified path reuse it. Explicit PRODUCER_EXPERIMENTAL_FAST_CAPTURE=true keeps old attempt-and-gate behavior. - Blank-frame retry can no longer splice wrong-frame pixels: recapture goes through recaptureDrawElementFrameForVerify — no static-dedup shortcut (lastEncodeResult runs ahead of the drain) and no "No cached paint record" screenshot fallback (post-injection that captures the canvas = the LAST drawn frame); any recapture failure falls back the whole render. - Verify indices derive from the producer-resolved duration (CaptureOptions.compositionDurationSeconds) instead of raw __hf.duration, so samples always land inside the drained range. - The platform clamp accepts "auto" GPU mode — the stock CLI resolves auto, and the literal-"hardware" clamp made default-on a no-op for the primary audience (masked in validation by explicitly-set env). Also: NaN-safe env parses (HF_DE_VERIFY / HF_DE_VERIFY_MIN_DB / HF_DE_BATCH); video comps skip verification when the session has no frame injector (probe sessions — black-video truth false-positived); psnr infrastructure failures skip the sample instead of failing the render; boundary-saturated sample indices are skipped; shader-transition comps prefer page-side compositing over default drawElement and compile-gated comps get page-side compositing restored; observability.clearFailure un-brands the recovered first streaming attempt; canary suite exempts known-marginal "any" comps from the cross-path PSNR gate; dead we-render options removed; clamp tests pin their env. Validated: canary suite 7/7; auto-GPU bare render engages the full stack; disk-path and worker-encode-off renders disengage default drawElement; malformed HF_DE_VERIFY_MIN_DB still verifies at the default threshold; engine suite 890 passed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(engine,producer): review fixes — PSC intent, verify-threshold clamp, fail-closed canaries Addresses miguel-heygen's review on #1998: - Page-side compositing restore preserves explicit caller intent (blocker): resolveConfig now records pageSideCompositingAutoDisabled only when IT turned page-side compositing off because drawElement was on; the compile-time drawElement gates restore page-side compositing only when that flag is set. An explicit enablePageSideCompositing:false from the programmatic API or HF_PAGE_SIDE_COMPOSITING=false stays off. Pinned by two config tests. - HF_DE_VERIFY_MIN_DB clamped to [10, 60] with a warning on out-of-range values: below ~10dB the check passes severe damage; above ~60dB natural encoder differences force a screenshot fallback on every verified render. - de-canary-suite.sh + de-gatecheck.sh run under set -euo pipefail with explicit `|| true` on expected-nonzero commands (render exits handled by the suite's own checks, grep no-match, kill/pkill/wait races) and a hard FAIL when the PSNR compare produces no value — release canaries fail closed. Full suite re-run green (7/7) under the new flags. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d5ecb013d7 |
perf(engine,producer): batch N drawElement frames per CDP round-trip (HF_DE_BATCH) (#1928)
Amortizes per-frame CDP protocol overhead (~3.5-9ms/frame) by looping seek -> paint-wait -> drawElementImage -> createImageBitmap in ONE page.evaluate for runs of consecutive frames; bitmaps still post to the encode worker per frame. Validated on 19 stratified DE comps: median 1.20x on top of worker-encode (to 1.56x), zero damaged frames, edge comps (static-dedup-heavy, clip-cut) bit-identical; mid-batch failure re-captures via the per-frame path (screenshot-fallback semantics preserved). Off by default; opt in with HF_DE_BATCH=4. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1d0dbcd3b2 |
feat(producer): fast-capture render stages + remote bg-image localizer (#1920)
* feat(engine): drawElementImage capture service * feat(engine): 3D projection + compositor-effect risk gate * fix(engine): gate filter drop-shadow wherever blur gates (review) detectCssEffectRisk documented drop-shadow as a ~29dB damage case but only detected blur( in its three scan paths — a drop-shadow comp stayed on the fast path despite the gate's own correctness contract. Detect drop-shadow( in computed styles, stylesheet rules, and GSAP tween vars, pinned by a focused test that runs the real page-side closure against a DOM shim (computed / stylesheet / tween coverage + blur regression + effect-free null). Addresses miguel-heygen's blocker on #1918. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(engine): frame-capture core — fast-capture routing, worker-encode, dedup extension # Conflicts: # packages/engine/src/services/screenshotService.ts * fix(engine): document HF_FORCE_DRAWELEMENT as diagnostic-only; make armStaticDedup idempotent (review) Addresses miguel-heygen's blockers on #1919: - HF_FORCE_DRAWELEMENT promoted from a stale "SCRATCH/Uncommitted" comment to a documented diagnostic flag: it exists for upstream-Chromium repro work (gate-vs-API isolation, crbug 521861819 149-vs-151) and R&D on gated effect classes; renders under it may be damaged BY DESIGN since it bypasses gates whose thresholds encode measured damage. Never production; the safety-net blank guard also stands down under it so diagnostic frames arrive unmodified. - armStaticDedup is now idempotent: the drawElement init path arms dedup before canvas injection, then initializeSession called it again — the second run overwrote the armed state with skipReason="capture_mode" (captureMode is "drawelement" by then), producing contradictory telemetry (armed frames + a skip reason), and re-ran the verification seeks on the fallback path. It now no-ops once staticFrames or a skip decision exists. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(producer): fast-capture render stages + remote bg-image localizer --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1a3330a972 |
fix(compiler): don't scan media tags inside comments/scripts; DOM-check auto-start (#1938) (#1940)
The timing compiler scanned raw HTML with tag regexes that weren't
comment-aware, so a comment or script merely mentioning `<video>`/`<audio>`
was rewritten as a real element — injecting id/data-start/data-hf-auto-start
into the comment text. That phantom attribute then tripped the probe stage's
substring check (`html.includes("data-hf-auto-start")`), launching an
unnecessary browser probe on every render with an unexplained empty reasons list.
- Mask comments, <script>, and <style> regions before the tag scan, then
restore them verbatim (compileTimingAttrs, extractResolvedMedia).
- Replace the probe's substring match with a DOM query
(video[data-hf-auto-start]) and add "auto-start video(s)" to the reasons list.
|
||
|
|
bb066077b4 |
fix(producer): avoid reviving hidden DOM in HDR layers (#1935)
* fix(producer): avoid reviving hidden DOM in HDR layers * fix(producer): filter transition HDR DOM masks * fix(producer): keep hidden timed descendants masked |
||
|
|
34590649a0 |
perf(engine): extraction cache on by default with atomic publish and LRU gc (#1901)
* perf(engine): extraction cache on by default with atomic publish and LRU gc Warm re-renders now skip source-video frame extraction entirely (video_extract 400ms -> 13ms on a 4-video composition; outputs are pixel-identical, PSNR inf). What made default-on safe: - Atomic entry publish: frames extract into a unique .partial-<pid>-<uuid> dir, the completion sentinel is written there, and the dir is renamed into the final key atomically. Concurrent renders sharing a cache can duplicate work but can never serve a torn entry (previously documented as single-writer only). - Size-capped LRU gc: best-effort sweep after extraction evicts oldest-used entries past a 2 GiB default budget (HYPERFRAMES_EXTRACT_CACHE_MAX_MB) and clears crashed writers' partials. Entries younger than 60 min are never evicted so live renders keep their frames. - Default cache dir: <tmpdir>/hyperframes-extract-cache-<uid>. Opt out with HYPERFRAMES_EXTRACT_CACHE_DIR=off (or none/false/0); a non-writable dir degrades to uncached with a single warning instead of failing the render. * fix(engine): harden extraction cache publish and surface cache ops signals Review hardening for the default-on extraction cache: - Bypass the cache for HDR-converted intermediates: the key snapshot describes the original source, so publishing converted frames under it would poison later plain-SDR renders of the same trim. (The follow-up transform-keyed change re-enables caching for these.) - publishCacheEntry TOCTOU: adopt a concurrent writer's completed entry both before removing an apparently-stale dir and after a failed retry rename, so a winner's publish is never destroyed or reported as a failure. - Observability for the failure paths: cachePublishFailures, cacheGcEvictions, cacheGcBytesFreed, and cacheAgedPartialsCleared on ExtractionPhaseBreakdown; gcExtractionCache now returns sweep stats. * fix(engine): sweep superseded cache generations in gc After a SCHEMA_PREFIX bump, old-generation entries (hfcache-v2-*) no longer matched the sweep's prefix filter and would orphan their disk forever. The gc now matches any hfcache-v* generation; superseded entries never receive sentinel touches, so the LRU evicts them first. |
||
|
|
7e8a1466c3 |
fix: producer render diverges from preview for sub-composition root styling (#1886)
Fixes #1847 The producer's render path stripped a sub-composition's authored root element and inlined only its children, so any CSS anchored on that root (its id or classes) matched nothing in the compiled HTML even though it resolved fine in Studio preview. Changes: - Wire flattenInnerRoot into the producer's sub-composition inliner (packages/producer/src/services/htmlCompiler.ts) so its render-time DOM shape matches the preview bundler's. - Rewrite a bare root [data-composition-id="X"] box selector to a :has()/:not() pair that lands on exactly one of the host or the flattened wrapper (packages/core/src/compiler/compositionScoping.ts), avoiding double-applying additive properties like padding. - Restore the composition's own id onto the flattened wrapper when the host has no id of its own, an "anonymous" host (packages/core/src/compiler/inlineSubCompositions.ts). - Fix the runtime's startResolver to find a composition's start time through the post-inlining data-composition-file marker, not just data-composition-src or data-composition-id (packages/core/src/runtime/startResolver.ts). Also adds regression coverage for the literal issue #1847 repro (a class, not just an id, on the authored root, styled via a descendant selector), a test proving the runtime compositionLoader's anonymous-host path doesn't share this bug, and fixes stale test documentation and a misattributed code comment surfaced during review. Verified: 29-fixture Docker regression sweep on linux/amd64 (matching CI) run 3x clean, 967/967 core unit tests, full CI green. |
||
|
|
f40dbd86cf |
fix(producer): surface the reason when audio mixing fails instead of silently shipping video-only (#1854)
At least 4 independent post-release feedback reports of a render completing successfully (exit 0) with audio elements correctly authored and detected at compile time (audioCount > 0), but the final MP4 having no audio track — discovered only via ffprobe or manual playback, with the CLI giving no indication anything went wrong. Users worked around it by muxing the generated audio in manually with ffmpeg. Root cause: runAudioStage sets hasAudio from processCompositionAudio's success flag, but discarded its error field — the actual reason a per-element audio prep step or the final mix failed (source not found, extract failed, ffmpeg error) was computed and then thrown away. A real audio-mix failure was therefore indistinguishable from "no audio was authored": both just produced hasAudio: false with zero diagnostic output. Thread the mixer's error through as audioError (only set when audios.length > 0 but the mix failed) and log.warn it from both call sites (the main render path in renderOrchestrator.ts and the distributed plan() path) so a real failure is loud instead of silently downgrading to a video-only render. Tests: 4 new cases for runAudioStage (mixer error surfaced, generic fallback message when the mixer doesn't provide one, no audioError on success, no audioError when there's no audio to mix). renderOrchestrator.test.ts (68 tests) unaffected. plan.test.ts's one failure (an audio-bearing planHash determinism test timing out at 30s) is pre-existing — reproduces identically on unmodified main with these changes stashed. |
||
|
|
e10e61ceb2 | fix(producer): normalize system-primary font stacks (#1857) | ||
|
|
733f88cb1f |
feat(producer,cli): render-reliability telemetry counters for capture hardening (#1850)
Follow-up to the render-reliability batch (#1841/#1842/#1843). Threads two capture-reliability counters through the existing observability → CLI-telemetry pipeline (no new PostHog wiring) so #1842's hardening is measurable on dashboard 1783183: - transient-retry burn (CaptureAttemptSummary.reason gains "transient-retry"; counted into RenderCaptureObservability.transientRetries on BOTH the recovered and the still-failed paths via a shared helper). - OOM classification (memoryExhaustionDetected set when describeMemoryExhaustion classifies the failure). Surfaced as capture_transient_retries + capture_memory_exhaustion_detected render-event props. Tests cover the attempt tagging and the payload mapping. Further follow-up (different subsystems): encoder-frame-0-exit signal, and P1-3 pre-flight-rejection / P1-4 cli_env_check counters. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
6be46813a2 |
fix(render): pre-flight aspect-ratio / alpha preset mismatch with actionable guidance (#1843)
Users pick an --resolution preset whose orientation/aspect ratio (or alpha/HDR mode) conflicts with the composition; the render fails deep in the compiler with a cryptic message. ~8K err / ~1K users. - New shared pure helper checkOutputResolutionCompatibility in @hyperframes/parsers — single source of truth for aspect/alpha/HDR/downsample/non-integer-scale constraints; suggests the matching-orientation, tier-preserving preset. - CLI render pre-flight aborts early (before browser/ffmpeg) with an actionable, fix-suggesting message; resolveDeviceScaleFactor delegates to the same helper for identical defense-in-depth messages. - Suggest (not auto-select); defers when dims can't be determined rather than guessing. - suggestMatchingPreset keys tier off the -4k suffix so square-family swaps (square + landscape-4k -> square-4k) aren't downgraded to HD. - render.js DOM polyfill made a lazy import; render.test cold-import beforeAll hooks given a 30s timeout to absorb CI contention. Render-reliability workstream P1-3. Success measured on PostHog dashboard 1783183. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
c0c3abf0f1 |
fix(producer): harden capture against timeouts, transient tab deaths, and OOM (#1842)
Four independent capture-infra hardening changes for the P2-5 failure bucket (~15K err / ~7K users): - protocolTimeout auto-scales by device-scaled output area (applied before probe launch, since it's immutable post ppt.launch()). - Single bounded transient retry (MAX_TRANSIENT_CAPTURE_RETRIES=1) on Target closed / Page crashed in the parallel disk-capture path; abort short-circuits before retry. - Narrow OOM classification (Set maximum size exceeded etc., disjoint from transient) → actionable guidance naming output dims. - StreamingEncoder.getExitError() threads FFmpeg's real exit reason into frame-0 encoder-death errors. Render-reliability workstream P2-5. Success measured on PostHog dashboard 1783183. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
cf573f7f3f |
fix(core,producer,cli): pre-flight validation for empty/malformed sub-compositions (#1831)
* fix(core,producer,cli): pre-flight validation for empty/malformed sub-compositions The #1 render failure bucket in production telemetry (PostHog project 356858, dashboard 1783183 "HyperFrames — Bottom-Line & Activation"; ~65-69K occurrences / ~27-28K affected users over 30 days, ~80% via AI-agent authoring flows) is a `data-composition-src` reference pointing at a scene file that is empty, malformed, or missing. Root cause, traced end-to-end: - The literal error "Composition HTML is empty or could not be parsed: <path>" is real (not a PostHog paraphrase) — thrown by a since-reverted guard in packages/core/src/compiler/inlineSubCompositions.ts (#1364), then changed to a silent skip in #1678 to avoid aborting renders on partial content during authoring. #1629 added per-assembler guards for 3 skill workflows (product-launch-video, faceless-explainer, pr-to-video), but general-video and hand-authored flows — where the dominant filename `scene-title.html` (40K+/68K of the bucket) originates — have no assembler and thus no guard. #1678 assumed the assembler guards from #1629 covered this pre-render; they only covered 3 of the many authoring flows. - On current `main`, an empty/malformed data-composition-src file no longer crashes or throws during render — it's silently dropped by the tolerant inliner. Reproduced locally: `hyperframes render` on a project with an empty scene-title.html "succeeds" after ~93s (two 45s pollSubCompositionTimelines timeouts) with the scene silently missing from the output video. `hyperframes validate` also reports "No console errors" for the same broken project. - The raw `Cannot destructure property 'firstElementChild' of 'documentElement' as it is null` crash reproduces directly against linkedom (the DOMParser polyfill packages/cli/src/utils/dom.ts installs in the real CLI runtime) for empty and non-HTML input — confirmed with a standalone repro script, not just inferred. jsdom/happy-dom (used in this repo's own test environment) are spec-compliant and never produce a null documentElement, which is why this needed a linkedom-specific test file. Fix: - New shared helper `checkSubCompositionUsability` (packages/core/src/compiler/subCompositionValidity.ts) is the single source of truth for "is this data-composition-src file usable" — mirrors the inliner's own parse/template/body logic so all callers agree. - `inlineSubCompositions.ts` (preview/studio bundling) now uses the shared helper internally but keeps its #1678 tolerant skip-and-continue behavior unchanged — mid-authoring iteration on a partial project must keep working. `onMissingComposition` now also receives a human-readable reason. - New render-only pre-flight (`assertSubCompositionsUsable` in packages/producer/src/services/htmlCompiler.ts) walks every data-composition-src reference (including nested ones, root-relative, matching parseSubCompositions' own resolution) before any compilation work starts, and throws naming every offending file at once. This is unconditional — not gated behind --strict — because a render that silently drops a scene is strictly worse than one that refuses to start. Confirmed locally: render now fails in ~0.4s with an actionable message instead of "succeeding" after 93s with a missing scene. - New `hyperframes lint` rule `missing_or_empty_sub_composition` (packages/cli/src/utils/lintProject.ts) surfaces the same check as a file-scoped, actionable lint error (already unconditional — lint exits 1 on any error). - `hyperframes validate` now also runs this check before launching a browser, so it no longer reports "No console errors" for a project with a broken sub-composition. - `packages/core/src/parsers/htmlParser.ts`: guarded every `documentElement`-may-be-null access (parseHtml, updateElementInHtml, addElementToHtml, removeElementFromHtml, extractCompositionMetadata, validateCompositionHtml) with a new typed `CompositionHtmlParseError` (or, for validateCompositionHtml's collect-and-report contract, a typed validation failure) instead of a raw crash. Tests: empty file, whitespace-only, malformed/non-HTML, missing file, nested sub-compositions (both happy path and broken-grandchild), and the happy path — at the shared-helper, lint, and render pre-flight layers. Not changed: the AI-agent authoring skills (skills/*). general-video and hand-authored flows have no assemble-index.mjs equivalent to guard, so the fix is at the CLI/render layer instead — flow-agnostic, covers every authoring path, and the skills' existing "run lint/validate and stop on failure" guidance now actually catches this class of mistake once run. Not run in this environment: the producer package's full regression-harness test suite (`bun test` in packages/producer) — it performs heavy real rendering (S3 asset downloads, Google Fonts fetches, full video encodes) and did not complete in a reasonable time in this sandbox. Verified instead via the targeted test file for all touched code (76/76 passing), whole-repo typecheck/build/oxlint, `fallow audit` (complexity/duplication/dead-code gate, clean), and manual end-to-end CLI runs (render/lint/validate) against reproduction projects, including a nested sub-composition scenario. CI should run the full producer suite before merge. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * refactor(parsers,lint): port empty-composition pre-flight to extracted packages Rebased onto main, which extracted @hyperframes/lint from core (lint depends only on parsers, not core). Relocate checkSubCompositionUsability from core to @hyperframes/parsers so both core (inliner) and lint can consume it without a core<->lint cycle; core keeps a @deprecated re-export shim. Correctness fixes from code review: - checkSubCompositionUsability now returns "no-composition-root" when the <template>/<body> content has no [data-composition-id] element (previously a marker-free placeholder body passed both guards). - lint's missing/empty sub-composition rule now only checks files reachable via data-composition-src from the root (matching render pre-flight), instead of a raw filesystem walk that false-positived on orphaned files. - drop `as string` cast in inlineSubCompositions in favor of an explicit null guard (per CLAUDE.md). Review-comment items: - move EmptyCompositionError JSDoc above the class (was above the adapter fn). - correct stale circular-ref comment to match actual silent-skip behavior. - rewrite self-contradicting lint message ("silently drop") to describe the new loud render-pre-flight abort. - add the __PLACEHOLDER__ (/^__[A-Z_]+__$/) skip to the render pre-flight so it agrees with lint. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
b33d54f54b |
fix(engine): retry probe on pollHfReady zero-duration timeout (#1824)
Renders were failing outright with "[FrameCapture] Composition has zero duration. Runtime ready: false, ..." whenever window.__renderReady didn't flip true within playerReadyTimeout (45s) — most often under host contention (e.g. several renders running concurrently), never from a defect in the composition itself. Confirmed by re-running an affected composition standalone: it succeeded immediately (initMs ~3.5-4.4s vs. the 45s timeout it hit under concurrent load). The probe stage already retries once with a fresh browser session for exactly this class of "succeeds on retry" infra flakiness (frame detachment, disconnects, navigation timeouts, launch failures), but isTransientBrowserError didn't recognize this message, so it fell through to an immediate, unretried failure. Match "Composition has zero duration ... Runtime ready: false" as transient. Left the "Runtime ready: true" case (pollHfReady's fast-fail: no GSAP timeline and no data-duration) unmatched — that's a genuine authoring bug, not a timing fluke, and should keep failing fast. |
||
|
|
0dfedd111c |
fix(core): root-cause id-less media wash in timingCompiler getAttr, drop the band-aid (#1792)
* fix(core): root-cause the id-less media wash in getAttr, drop the band-aid The blank-wash/dropped-audio fix in #1790 added assignMissingMediaIds in the producer to stamp ids onto id-less timed media. That was a band-aid: the real cause is timingCompiler's getAttr, whose regex had no name boundary at all, so getAttr(tag, "id") matched the trailing id="…" inside data-hf-id="…". compileTag saw a phantom id and skipped its existing hf-video-N/hf-audio-N injection, leaving the element with no real el.id — which the render pipeline keys off of. Fix getAttr with the same (?<![\w-]) lookbehind used for the lint readAttr fix. compileTag's auto-id injection now fires for data-hf-id-only media, in both the main composition and sub-compositions (parseSubCompositions runs the same compileTimingAttrs pass), so assignMissingMediaIds is removed entirely. Extends the regression fixture with a standalone id-less <audio> (the dropped- audio side, previously untested) and raises minAudioCorrelation to 0.9. Adds a timingCompiler test for the data-hf-id/id boundary. * test(producer): use seeded pink noise (not a pure sine) for fixture audio A continuous sine anti-aligns under the audio cross-correlation (correlation -1.0 from a sub-period offset). Broadband seeded noise correlates robustly. * test: cover audio-side id injection via unit test; keep render fixture video-only The audio render-baseline used synthetic sine/noise, which anti-aligns under the harness audio cross-correlation (deterministic -1.0). Real audio fixtures are unaffected. Cover the audio side of the boundary fix with a deterministic timingCompiler unit test (id-less <audio> gets hf-audio-N) instead, and keep the render fixture video-only. * test(producer): regenerate baseline under the root fix (hf-video-N from compileTag) |
||
|
|
74f9c31b3f |
fix(producer,lint): id-less media renders blank wash instead of footage (#1790)
* fix(producer,lint): id-less media renders blank wash instead of footage A timed <video>/<audio> identified only by a Studio-stamped `data-hf-id` (no real `id`) rendered as a flat white/grey wash with dropped audio, and lint stayed silent so it surfaced only at render. Root cause, two layers: - lint `readAttr(tag, "id")` used a `\b` boundary, which treats the hyphen in `data-hf-id="…"` as a word break — so reading "id" matched the trailing `id="…"` inside `data-hf-id` and returned a phantom id. `media_missing_id` therefore never fired for media carrying only a data-hf-id. Switched to a `(?<![\w-])` lookbehind so a short name can't match the tail of a longer hyphenated attribute (also fixes "width" matching `data-width`, etc.). - the render pipeline identifies media by the real `el.id`: frame extraction keys injected stills as `__render_frame_<id>__`, the runtime frame-swap matches on `el.id`, and the audio mixer selects `audio[id][src]`. An empty `el.id` meant injected frames/audio never matched. compileForRender now assigns a stable positional id to every id-less timed media element before any stage parses or serves the HTML. Adds a producer regression fixture (video with data-hf-id, no id) and a lint test covering the data-hf-id/id collision. Baseline mp4 generated separately. * test(producer): baseline for video-hfid-no-id regression fixture Golden compiled.html + output.mp4 (generated on linux/amd64 in the Dockerfile.test image). Compare-mode passes: compilation, visual (0 failed frames), and audio (correlation 1.000). A regression to the blank-wash behaviour fails the visual check. |
||
|
|
c9613cd826 | fix(producer): keep video captures viewport-bound on software (#1788) | ||
|
|
35a01d9058 |
perf(engine): reduce init overhead in headless capture sessions (#1718)
Flush the GSAP proxy queue synchronously during capture session initialization and parallelize independent media/font/tailwind readiness waits. Closes #1715. Co-authored-by: Miguel Angel Simon Sierra <miguel.sierra_miga@heygen.com> |
||
|
|
7a4853dfe6 |
refactor: extract @hyperframes/studio-server from core (#1757)
* refactor: extract @hyperframes/studio-server package from core Moves all studio-api routes, helpers, and Hono server wiring from packages/core/src/studio-api/ into a new standalone packages/studio-server package (@hyperframes/studio-server). Core keeps thin re-export stubs at @hyperframes/core/studio-api and the subpath helpers (screenshot-clip, draft-markers, etc.) for backward compatibility. Consumer imports (cli studioServer, vite adapter/config, producer htmlCompiler, studio manualEditsTypes) are updated to import from @hyperframes/studio-server directly. Also exports rewriteInlineStyleAssetUrls from @hyperframes/core root (was in compiler/rewriteSubCompPaths.ts but not re-exported), required by @hyperframes/studio-server/helpers/subComposition. Removes postcss-selector-parser from @hyperframes/core dependencies (moved to @hyperframes/studio-server which owns the routes that used it). Depends on @hyperframes/parsers (PR #1755). * fix(ci): add parsers+studio-server to Dockerfile and build before preview tests * fix(ci): build @hyperframes/studio-server before Test and studio load smoke Studio's vite.config.ts imports @hyperframes/studio-server, which resolves via its "node" export condition to built dist. The Test and studio-load-smoke jobs only built parsers + core, so esbuild's config load failed to resolve the package entry. Build studio-server too. * fix(studio): repoint sdkCutoverParity test import to studio-server sourceMutation moved from core's studio-api to @hyperframes/studio-server; the test still imported the deleted core path. This was masked while studio's vite.config failed to load (couldn't resolve studio-server); now that the config loads, the test runs and the stale import surfaced. |
||
|
|
98d0bdd73c |
refactor: extract @hyperframes/lint from core (#1756)
* refactor: extract @hyperframes/lint package from core Moves all lint rules, hyperframeLinter, lintProject, and related types from packages/core/src/lint/ into a new standalone packages/lint package. Core keeps a thin re-export stub at @hyperframes/core/lint for backward compatibility. Consumer imports (cli lint command, producer hyperframeLint) are updated to import from @hyperframes/lint directly. Depends on @hyperframes/parsers (PR #1755). * fix: restore postcss-selector-parser in core (sourceMutation.ts still uses it) * fix(ci): add parsers+lint to Dockerfile and build before preview tests * chore: update bun.lock after restoring postcss-selector-parser dep * test(cli): update lintProject test for string-dir signature from @hyperframes/lint * refactor(core): single-source the lint engine in @hyperframes/lint Delete core's byte-identical copy of the lint rule engine and re-point staticGuard at @hyperframes/lint, so the render-time render-gate and the studio preview share one rule engine instead of two copies that could silently diverge. Back-compat preserved via the @hyperframes/core/lint stub. Addresses review feedback on the dual-copy footgun. |
||
|
|
c9e8dd3862 | fix(runtime): honor render fps when seeking (#1739) | ||
|
|
ca9e1316be |
perf(producer): stream binary file responses, async-read HTML (#1735)
* perf(producer): stream binary file responses, async-read HTML
Replaces the per-request readFileSync in fileServer's static file handler
with a createReadStream pipe (binary) and an async readFile (HTML). Static
asset serving no longer blocks the Node event loop.
Why
---
The pre-fix handler called readFileSync(filePath) on every binary asset.
On video-heavy compositions Chrome requests several 32MB video files
back-to-back; each readFileSync(32MB) blocked the main event loop long
enough to wedge concurrent /health responses and other timers.
Scope clarification — this addresses the event-loop block documented at
renderOrchestrator.ts:1277-1306 (the video-heavy regression class). It is
NOT the fix for today's infinite-duration incident; Miguel is shipping
that upstream as a plan()-time duration guard. The two are complementary:
- Miguel's guard kills the impossible-work input shape before chunk
planning so the producer doesn't try to enumerate 300B frames.
- This streaming fix removes the next-largest known main-thread block
(large binary I/O during video-heavy renders), so future wedge
classes don't kill otherwise-healthy probes either.
The companion worker_thread /health PR + the heygen-com/app probe-timeout
bump round out the defense-in-depth: even if some future code path
introduces another main-thread stall, the probe lives off-thread and the
budget is 30s anyway.
What changed
------------
fileServer.ts: switched both file branches off the sync I/O path.
- Binary (the hot path for video-heavy renders): readFileSync(filePath)
-> createReadStream + Readable.toWeb -> Response stream body.
Content-Length is set via statSync so Chrome's range-aware media
stack sees the size up front. The handler is now async because the
HTML branch awaits.
- HTML (small files; injected with pre/head/body scripts):
readFileSync(filePath, "utf-8") -> readFile(filePath, "utf-8").
The injection is still sync — pure string ops — only the disk read
moved off-thread. Index HTMLs are tiny (~200KB max for AI-generated
compositions) but a ms of stall per render-start adds up across a
fleet.
Test
----
fileServer.test.ts: added a streaming regression that pins three
properties on a 5MB synthetic binary asset (chunk-boundary spanning):
1. Correctness — served bytes match the file across multiple
createReadStream chunks (default 64KB highWaterMark).
2. Content-Length header is set from statSync.
3. Four parallel fetches all return identical content; the streaming
path doesn't serialize them.
All 31 fileServer tests pass locally (bun test).
TODO: link Miguel's upstream plan() duration guard PR once known.
— Jerrai
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(producer): implement Accept-Ranges + 206 Partial Content for fileServer
Delivers the range-request semantics the original PR body promised but
the diff did not implement. Without range support, Chrome's <video>
element issues full-file GETs on seek; with this commit it can issue
`Range: bytes=...` and get a sliced 206 back, so seek + partial-load
work without re-pulling the whole file.
- Add `parseRangeHeader` (exported for unit tests) covering the three
RFC 7233 single-range forms: bytes=START-END (closed), bytes=START-
(open-ended), bytes=-SUFFIX (last N bytes). Multi-range falls back to
`absent` (full 200) so we never reassemble multipart/byteranges.
- Binary path now returns 206 Partial Content with Content-Range +
sliced Content-Length on satisfiable ranges, 416 Range Not Satisfiable
with `Content-Range: bytes (asterisk)/<size>` on unsatisfiable ranges,
and 200 with `Accept-Ranges: bytes` on full-body GETs so clients know
ranges are supported.
- Add unit tests for parseRangeHeader (10 cases: 3 forms, clamping,
unsatisfiable edges, malformed inputs, multi-range fallback).
- Add integration test covering 200 + Accept-Ranges, all 3 range forms
with byte-correct slices, 416 on out-of-bounds, and multi-range -> 200
fallback.
Addresses Miga's review finding on #1735.
Co-Authored-By: Jerrai <noreply@anthropic.com>
— Jerrai
---------
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
01a10cdc53 |
fix(producer): serve /health from a worker_thread so probes survive main-thread stalls (#1733)
* fix(producer): serve /health from a worker_thread so probes survive main-thread stalls Adds an off-main-thread /health endpoint that listens on its own port (default 9848, env PRODUCER_HEALTH_PORT). The endpoint binds inside a Node worker_thread with a minimal node:http server — separate event loop, separate isolate — so probe responses don't depend on whatever the producer's main thread is doing. Why now ------- Today's hyperframes-producer crashloop traced to an infinite GSAP timeline -> distributed planner trying to enumerate ~300,000,000,000 frames -> sidecar /health stops landing within k8s's 5s window -> otherwise-healthy pods killed. Miguel is shipping the root-cause fix at plan() time (impossible / non-finite / sentinel durations get rejected before chunk planning). That removes today's wedge. This change is defense-in-depth for the kill mechanism. Even with the plan() guard, future wedge classes can stall the main event loop for seconds at a time: large synchronous file I/O (see the companion fileServer streaming PR), GC pauses on long-running renders, tight loops in user-authored GSAP / Three.js / canvas code, future activity / pool changes whose runtime cost we haven't yet characterized. Probe responsiveness should reflect process liveness, not main-thread event-loop responsiveness. If the entire Node process is dead the OS tears down both threads' sockets simultaneously and k8s correctly kills the pod. Anything short of that and the worker thread's listener keeps answering. Backwards-compatible: the main-thread /health on PRODUCER_PORT (9847) keeps working exactly as before. The k8s sidecar probe config in heygen-com/app can migrate to the worker port at its own pace. A companion heygen-com/app PR in this batch raises the probe timeout from 5s -> 30s as a last-resort backstop. TODO: link Miguel's upstream plan() duration guard PR once known. Test: healthWorker.test.ts (vitest) — 3 tests pass locally, including the load-bearing one: stays responsive while the main thread is blocked on a 500ms sync busy-spin. — Jerrai Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(producer): tighten healthWorker startup race + shutdown semantics Addresses Miga's review on #1733. - server.ts: store the worker as a Promise<HealthWorkerHandle | null> instead of mutating a `let` from inside `.then`. A SIGTERM landing before the `.then` callback fired would previously see `healthWorker === null` and skip cleanup. shutdown() now `await`s the promise with a bounded 1.5s timeout so a hung-startup worker can't keep SIGTERM waiting (worker.terminate() from process exit still kills it). - healthWorkerThread.ts: replace `process.exit()` inside the worker with `parentPort.close()` + natural event-loop drain. Node-version semantics for `process.exit()` from a worker have been historically inconsistent; the documented clean path is to close the channel and let the worker exit naturally. Also drops the redundant 2s force-exit on shutdown — the parent already owns the authoritative deadline via Promise.race + worker.terminate(), so the worker-side timer was belt-and-suspenders noise. Co-Authored-By: Jerrai <noreply@anthropic.com> — Jerrai --------- Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
dca5fa90cb | fix(producer): reject impossible distributed durations (#1732) | ||
|
|
7db84fc0ad |
fix(producer): rebuildExtractedFramesFromPlanDir off-by-one in framePaths key indexing (#1730)
* fix(producer): rebuildExtractedFramesFromPlanDir off-by-one in framePaths key indexing In distributed chunk-lambda render mode, every <video>'s first-paint frame (the moment a vid first becomes visible on the composition timeline) renders as PRISTINE Y=16 black. For a 3-vid back-to-back composition (v1: 0-4s, v2: 4-8s, v3: 8-12s at 30fps), frames 0, 121, 242 are all PRISTINE black; the render then either stays black for 1 frame, or shows body bg + persistent overlays only (Y~22 with sparse highlights from text/logo). The symptom only reproduces in distributed mode — local single-process renders are unaffected. Root cause: rebuildExtractedFramesFromPlanDir builds the framePaths Map with 1-based keys, but the consumer (getFrameAtTime at engine/videoFrameExtractor.ts:958) computes a 0-based frame index via Math.floor(localTime * fps + 1e-9). For each vid's first-paint frame (localTime === 0 → frameIndex === 0), framePaths.get(0) returns undefined; the vid is silently dropped from activePayloads, videoFrameInjector doesn't fire, syncVideoFrameVisibility hides everything, and BeginFrame screenshots an empty composition. Every other site in the codebase builds/consumes framePaths with 0-based keys: - engine/videoFrameExtractor.ts:317 framePaths.set(index, ...) - engine/extractionCache.ts:204 framePaths.set(idx, ...) - engine/videoFrameExtractor.test.ts:264/1066 framePaths.set(i, ...) - engine/videoFrameExtractor.ts:958 (consumer) Math.floor 0-based - producer/renderOrchestrator.test.ts:287/315/349 framePaths.get(0) Only producer/distributed/renderChunk.ts:198 was 1-based, with a stale comment claiming FrameLookupTable indexes frames 1-based — which the surrounding evidence contradicts. This is why local tests pass while distributed-lambda renders always had cold black at each vid first paint. Verified locally against a 3-vid composition and a single-vid 4-worker case in a Lambda render fleet. Before fix: every vid first-paint frame is PRISTINE Y=16 black. After fix: all frames are valid source content, blackdetect reports zero black regions outside legitimate source video content (intentional fade-ins / hard cuts in source mp4). * test(producer): pin rebuildExtractedFramesFromPlanDir 0-based framePaths contract Regression guard for the off-by-one fix in HF#1730. The pre-fix code indexed framePaths 1-based while the consumer (getFrameAtTime in engine/videoFrameExtractor.ts:958) reads 0-based, dropping every <video>'s first-paint frame in distributed chunk-lambda renders. Asserts framePaths.get(0) resolves to the first extracted frame, and framePaths.get(N-1) resolves to the last — pre-fix the keys were shifted to [1..N], so get(0) returned undefined and get(N) resolved. Verified to fail against the previous i+1 indexing. Also exports rebuildExtractedFramesFromPlanDir (was module-local) so the test can call it directly. Pure logic worth testing in isolation — the bug only reproduces under distributed mode and the existing renderChunk.test.ts already pays a multi-second Chrome smoke probe in its module-level beforeAll, so the regression check lives in its own file (rebuildExtractedFrames.test.ts) and runs Chrome-free in ~10ms. The function's doc comment said "1-based framePaths" — updated to "0-based" with a pointer to the consumer site and the bug context. Per Miguel's REQUEST_CHANGES on HF#1730. — Jerrai (https://claude.com/claude-code) --------- Co-authored-by: James <james.russo@heygen.com> |
||
|
|
92385711dc |
fix(engine): hold last frame when a clip's media is shorter than its slot (#1726)
Renders showed the page background (a one-frame black flash) right before a cut when a video clip's source media was a hair shorter than its data-duration slot — the common case, since `ffmpeg -t 1.45` emits 43 frames = 1.433s at 30fps. The frame lookup only held the last frame at the exact clip end, so the sub-frame remainder rendered blank. - Hold the last extracted frame for the rest of the slot once the source is exhausted, within a tolerance floored at the compiler's 0.05s clamp epsilon so the seam is covered at any fps (2 frames alone is < 0.05s above 40fps). Clips deliberately much shorter than their slot still blank for the tail (unchanged). - Warn when the compiler clamps a video's data-duration down to its media length (slot longer than source by more than the clamp epsilon): a render-time `[compile]` warning in the producer, plus a matching `validate` warning that reads each <video>'s live duration in headless Chrome (static HTML lint can't see media durations). A shared `analyzeClipMediaFit` keeps both on one threshold. Adds engine unit tests for the hold behavior and the analyzer. |
||
|
|
690cf1b7a5 |
fix(producer): stop retrying capture attempts that made zero progress (#1725)
* fix(producer): stop retrying capture attempts that made zero progress A structurally broken composition (never-ready page, zero duration, or unparseable HTML) captures no frames, so the adaptive retry loop kept re-running it at halved parallelism — 16->8->4->2->1 workers — each attempt burning a full readiness/protocol timeout per worker. That multiplied wall-clock to ~46min on broken renders and was the driver of the render P95 blowup (~370k -> 2.79M ms) seen Jun 20-22. Add captureAttemptMadeProgress(): when an attempt leaves at least as many frames missing as it set out to capture, it made no forward progress, so the composition is broken rather than the workers being flaky. Bail immediately instead of retrying. A partially-captured attempt still retries, so genuine flaky-worker gaps are unaffected. * fix(producer): log the zero-progress bail + cover it with an integration test Address review feedback on the no-progress capture guard: - Warn before bailing so an oncall can tell a structurally-broken render that bailed fast apart from one that exhausted worker-halving retries (both previously threw the same "frame(s) are missing" message). - Add an integration test that drives executeDiskCaptureWithAdaptiveRetry through the bail (capture functions mocked to write nothing) and asserts a single attempt runs — the gate would otherwise walk 4->2->1 workers. Guards the placement of the gate, not just the predicate. - Reword the helper docstring (drop stray prefix and internal incident detail). |
||
|
|
0558b8761e | fix(producer): retry probe navigation timeouts (#1713) | ||
|
|
89ff299a11 |
fix(engine): defend macOS regular Chrome screenshots
Fixes #1699. |