mirror of
https://github.com/heygen-com/hyperframes.git
synced 2026-09-05 10:14:30 +00:00
f44bc3a5253af417705b9a373a9f0338a58dae08
213
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f44bc3a525 |
fix(engine,producer): drive forceScreenshot from one authoritative local
Miguel R4 blocker on #2359: my R3 fix at renderOrchestrator only updated the observability copy, leaving the authoritative captureForceScreenshot local at compileResult.forceScreenshot (false for auto→software). The frameCapture side clamped its own local and correctly routed screenshot, but downstream orchestrator code overwrote observability back to beginframe from the still-false local at two sites: - Parallel-stream label at renderOrchestrator.ts:2293 mis-labelled the stream as 'beginframe' when actual capture was 'screenshot'. - capture_strategy telemetry at renderOrchestrator.ts:2440-2450 overwrote the earlier observability correction, so the final captureMode observation flipped back to 'beginframe' while the engine actually captured screenshot. Fix: extract the clamp into a caller-facing helper applyConcreteGpuScreenshotClamp(current, resolvedGpuMode, cfg) that returns the (possibly-promoted) new boolean. Callers assign it back to their authoritative local, so routing + telemetry + strategy code read one value. Changes: - packages/engine/src/config.ts: new exported applyConcreteGpuScreenshotClamp; delegates to shouldClampToScreenshotForConcreteGpu but computes the caller's final value, not just the clamp decision. Reads the programmatic opt-out from cfg.forceScreenshotExplicitlyOptedOut. Idempotent on already-true input. - packages/engine/src/index.ts: export the new helper. - packages/engine/src/services/frameCapture.ts: replace the inline OR expression with applyConcreteGpuScreenshotClamp. - packages/producer/src/services/renderOrchestrator.ts: assign result into the AUTHORITATIVE captureForceScreenshot local (was updating only observability). Downstream parallel-stream label at :2293 and capture_strategy telemetry at :2440-2450 now read the corrected value. Tests: 6 new caller-level cases for applyConcreteGpuScreenshotClamp covering the exact matrix Miguel called out: - resolved software + default false → promotes to true (screenshot) - resolved software + programmatic opt-out → stays false (BeginFrame) - resolved hardware + default false → stays false - resolved software + already-true → stays true (idempotent) - resolved software + env PRODUCER_FORCE_SCREENSHOT=false → stays false - resolved software + undefined cfg → promotes to true (frameCapture path) Local: 67/67 engine config tests pass (was 61). oxfmt clean. |
||
|
|
72daac2a1d |
fix(engine): carry programmatic forceScreenshot opt-out to concrete-resolved site
Miguel R3 blocker on #2359: the runtime helper only checked the env opt-out (PRODUCER_FORCE_SCREENSHOT=false), silently defeating the documented programmatic escape hatch (overrides.forceScreenshot === false) on the browserGpuMode:'auto' → software probe path. At the concrete-resolution site the boolean forceScreenshot === false is ambiguous between default and explicit opt-out — resolveConfig sees the provenance but the runtime helper does not. Fix: persist provenance on the resolved config. - New INTERNAL EngineConfig field forceScreenshotExplicitlyOptedOut, set by resolveConfig when EITHER env or programmatic explicit-false is present. Purpose-documented in the type as 'not intended to be set by callers'. - shouldClampToScreenshotForConcreteGpu gains an opts.programmaticOptOut parameter; returns false early when set. Env stays as the third arg (backward compatibility with existing tests). - frameCapture.ts and renderOrchestrator.ts pass config.forceScreenshotExplicitlyOptedOut through at both call sites, so the auto→software probe path preserves the same escape hatches as literal browserGpuMode:'software'. New tests: 5 additional cases across the helper (programmatic opt-out alone; programmatic beats missing env) and resolveConfig provenance (programmatic sets flag; env sets flag; neither leaves it undefined). Local: 61/61 engine config tests pass (was 56). |
||
|
|
2e44602f99 |
fix(engine,producer): apply software-GPU screenshot invariant at concrete-resolved point
Addresses Miguel's R1 blockers:
1. `browserGpuMode: "auto"` that runtime-probes to software slipped past the
`resolveConfig` clamp — that clamp only sees the pre-resolve string. Add
`shouldClampToScreenshotForConcreteGpu(resolvedGpuMode, currentForceScreenshot, env)`
in `packages/engine/src/config.ts` and apply it at BOTH concrete-resolution
sites:
- `packages/engine/src/services/frameCapture.ts`: downgrades `preMode`
from "beginframe" to "screenshot" when resolved GPU is software (respects
`PRODUCER_FORCE_SCREENSHOT=false` env opt-out), fixing the routing.
- `packages/producer/src/services/renderOrchestrator.ts`: updates
`captureObservability.forceScreenshot` (and thus `captureMode`) at the
same call site, fixing the observability truth on the auto → software
case.
2. New unit tests in `config.test.ts`:
- Documents the auto-branch gap (resolveConfig leaves auto as
forceScreenshot=false — the runtime companion closes it).
- 5 branch tests on `shouldClampToScreenshotForConcreteGpu` covering
software / hardware / already-forced / env-opt-out / non-"false" env
values.
Full suite: 56/56 pass.
Scope narrowing on Blocker 2: the distributed rendering path at
`packages/producer/src/services/distributed/plan.ts:753-754` and
`renderChunk.ts:462-466` explicitly hardcodes `browserGpuMode:"software",
forceScreenshot:false` post-resolveConfig and stays outside this PR's
invariant boundary. `compileStage` may still flip it to true for alpha
formats, but generic MP4 distributed renders on SwiftShader hosts remain
BeginFrame. That's a separate architectural cleanup (needs its own
behavior-change trace); the PR body now scopes the invariant to the
in-process CLI/orchestrator path.
|
||
|
|
3df59fc0a4 | fix(engine): support current FFmpeg filter scripts (#2324) | ||
|
|
fecd7dc1d3 |
Merge pull request #2248 from heygen-com/bf-reuse-telemetry
feat(producer): surface beginframe no-damage reuse counters in perf summary and telemetry |
||
|
|
1b95af8bbb | fix(engine): densify drawelement self-verify with parallel worker count | ||
|
|
dde3afb72e | feat(producer): surface beginframe no-damage reuse counters in perf summary and telemetry | ||
|
|
0809c98518 | fix(producer,engine): gate drawelement on ancestor background-image + tail verify sample | ||
|
|
a355fb2f6b |
fix(producer,engine,cli): oom wrapping, cancellation, fallback-reason gaps
Three defects found by max-effort code review of this branch: 1. The Bun OOM exact-match regex was defeated by this codebase's own parallel-worker error wrapping. executeParallelCapture/formatWorkerFailure (parallelCoordinator.ts) always wrap a worker's error as "Worker N: <message>", optionally suffixed and joined with other workers' segments, all prefixed "[Parallel] Capture failed: ". That wrapping defeated the exact-message check for exactly the cohort (deParallelRouter routed, N separate Chrome processes) the OOM-drops-to-1 fix targets — a real OOM there would retry at the SAME worker count instead of dropping to 1. Added a second pattern that recovers the signal by requiring "out of memory" appear as the WHOLE content of a "Worker N: ..." segment (bounded by end-of-string/"; "), preserving the same exact-match property (no bare substring match) while surviving the wrapping. Verified against the real wrapping logic, not a hand-typed guess at its shape. 2. shouldRetryViaPinnedFallback didn't exclude cancellation, so aborting a render mid-capture on the pinned router/inversion cohort would detour through spawning a fresh encoder/capture session before the outer catch's RenderCancelledError branch ended the render — delaying "stop" with a pointless resource spin-up/tear-down. Added an isCancellation param (checked first, before isVerifyError) using the same `err instanceof RenderCancelledError || abortSignal?.aborted` check the outer catch already uses. 3. deFallbackReason (this PR's new "oom"/"capture_error" values) was set locally but never mirrored into RenderCaptureObservability alongside deSelfVerifyFallback, so a render that fails AFTER a fallback attempt (perfSummary never built) was indistinguishable in render_error telemetry from one that never attempted any fallback — undercutting the "how often does the OOM retry fire on a render that still ultimately fails" question this branch exists to answer. Threaded through RenderCaptureObservability → RenderObservabilityTelemetryPayload → renderObservabilityTelemetryPayload, mirroring the existing deSelfVerifyFallback plumbing. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
b3f244a7e9 |
fix(engine): recognize Bun/JavaScriptCore's OOM message in isMemoryExhaustionError
Found while testing the previous commit's OOM-drops-to-1-worker fallback end-to-end: the producer's deployed runtime is Bun (JavaScriptCore), not Node (V8) — see packages/gcp-cloud-run/Dockerfile's `bun dist/server.js` entrypoint. All 7 MEMORY_EXHAUSTION_ERROR_PATTERNS are V8-specific allocation failure signatures; JSC's equivalent for the same single-oversized-allocation RangeErrors is the bare string "Out of memory" (verified against real Bun behavior), which none of them match. Without this, isMemoryExhaustionError returns false for genuine production OOM, so the memory-specific worker-count reduction just added would never actually engage where it's deployed — every OOM would fall through to the generic capture_error retry path instead. Matches the FULL (trimmed) message only, not merely a substring — same rationale as the existing V8 patterns' comment: "out of memory" also appears in benign WebGL/GPU console noise that must not trip this classifier. Verified end-to-end from a script inside the producer workspace (importing the real @hyperframes/engine source, not a stale globally-cached npm dist a script outside the workspace would otherwise resolve to): a genuine Bun RangeError from new Uint8Array(Number.MAX_SAFE_INTEGER) now correctly classifies as memory exhaustion and drives both resolveInversionRetryPlan and resolveParallelRouterRetryPlan down to workerCount=1 on retry. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
b9321b7489 |
fix(engine): descend into repeating nested timelines for call() detection
Address PR #2093 review feedback (Miga, Rames D Jusso): - The walker treated a repeating nested timeline (total > single) as an opaque interval and never descended into it, so a tl.call() living inside one would slip past hasTimelineCall detection entirely — the "any tl.call() disqualifies" claim wasn't quite literal. Now recurses for detection purposes even when the span is already opaque; the parent-level interval still dominates for frame-animated-marking, so this only widens what counts as "has a call()," never narrows the existing interval coverage. - Restored the totalDuration() vs duration() rationale comment that got dropped when the tl.call() detection comment was added above it. |
||
|
|
50c4a10234 |
fix(engine): disqualify static-frame dedup on any tl.call()
Real bug report: a mono count span driven by a GSAP tl.call() (a counter going "0 sur 0" -> "1 sur 1" at a later beat) rendered the LATER value baked in from frame 0 of an EARLIER, unrelated static-hold span, despite the dedup log reporting "verified". Root cause: computeStaticFrameSet's tween walker only tracks property tweens, so a call()-driven textContent mutation carries no tracked interval and the span around it looks fully static. verifyStaticFramesSafe does catch genuine drift WITHIN a run it's checking, but a call() is a one-shot side effect wired as both onComplete and onReverseComplete (GSAP has no separate "undo" — crossing it in either direction fires the SAME forward mutation). Verifying a LATER run forward-seeks past the call(), permanently mutating the live page; an EARLIER run already passed its own check before that happened, so nothing re-verifies it afterward. Real capture then starts on the same corrupted page and bakes the wrong value into the earlier span's reused buffer. No reliable way to tell a DOM-mutating call() from a harmless one (analytics ping, class toggle) without executing it, so this disqualifies the whole comp on ANY call() — conservative, costs some dedup perf on comps that use call() harmlessly, but correctness over speed. |
||
|
|
381887541f |
fix(engine,producer): fix quadratic dedup rescan, correct race justification
Address two max-effort code-review findings on PR #2056 not covered by the earlier review-gap commit: - captureFrameToBufferPipelined's static-dedup reuse branch never advanced session.lastEncodeResultFrame, unlike its sibling real-capture branches. The gap-check window is computed from that watermark, so every consecutive reuse in a static run rescanned an ever-widening window instead of just the newest frame — O(n^2) total work over a long static stretch instead of O(n). - The "single-threaded, no race" justification on the shared parallelGuard closure was wrong: the guard has real internal await points (recapture, PSNR) between reading and writing its sizes/absFloor/acceptedSmall state, so concurrent workers' calls do interleave there (confirmed). Replaced with the actual reason it's safe: absFloor only ratchets down, sizes is append-only and order-independent for the median, and acceptedSmall's fast path re-validates by exact byte-equality regardless of which worker wrote the reference buffer. |
||
|
|
2dbe958a49 |
fix(engine,producer): close review gaps in parallel drawElement streaming
Address PR #2056 review feedback: - Fix totalFrames progress inflation for interleaved tasks — divide each task's span by its frameStride to match the actual per-worker frame count (captureFrameRange steps by stride), instead of summing raw endFrame-startFrame which double(N)-counts interleaved tasks. - Attach a no-op .catch to each frame's pipelined encodeResult at kick time so an abandoned promise (loop exits early on abort/error before draining it) can't surface as an unhandled rejection during teardown. - Document why the pipelined branch's stride=1 path is validation-only in production (HF_DE_PARALLEL_STREAM always uses interleaved distribution) so a future refactor doesn't unknowingly widen it. - Comment the intentional single shared parallelGuard/parallelStats across workers (safe single-threaded, better rolling-median signal). |
||
|
|
b3493a7b61 |
feat(engine,producer): verified interleaved parallel drawElement streaming (opt-in)
Step 2 of the DE engagement plan: multi-worker drawElement capture through the streaming encoder, with the full runtime self-verification net riding along — the confinement rule that kept the parallel clamp in place is now satisfied on this path. Opt-in via HF_DE_PARALLEL_STREAM=true; default routing (including the #2026 single-worker inversion) is unchanged. Mechanism: - distributeFramesInterleaved + WorkerTask.frameStride: worker i captures frames i, i+N, i+2N... — seek-based capture makes stride free and the ordered writer's reorder window shrinks from totalFrames/N to N (contiguous chunks serialize workers behind the writer). - Depth-2 pipelined worker-encode produce in the parallel worker loop (the same shape as the sequential loop; frame k's in-page encode overlaps k+stride's produce). HF_DE_PAR_DEBUG=1 traces the first frames per worker. - Drain guard extracted to createDrainFrameGuard (session-parameterized): every parallel frame gets the SAME blank-guard + PSNR self-verify as the sequential drain, against its owning worker's pre-injection ground truth (all sessions arm identical sample indices from CaptureOptions.compositionDurationSeconds). - FrameReorderBuffer.abort(err): a failed worker (e.g. verification error) rejects all parked and future waiters — without this, peers park forever in waitForFrame and the pool (which awaits ALL workers before surfacing errors) deadlocks. Found by the verify-trip test; unit-tested. - The typed DrawElementVerificationError is preserved past the pool's error-string flattening so the orchestrator's verify-retry recognizes it. - Static-dedup stride hazard fixed: lastEncodeResult reuse now requires EVERY frame in (lastEncodeResultFrame, i] to be predicted-static (sequential capture reduces to the old has(i) check). - Workers get separate browser PROCESSES under the flag: pages co-tenant in one browser starve non-active pages of BeginFrames on the paint-wait path (measured 86s vs 30s on a 3,245-frame rAF comp). Validation: - Happy path W3: verify samples pass across workers (4x inf on the 2,381f comp), output vs single-worker DE = 59.3dB (encode noise floor) — the interleave + dedup-stride produce identical pixels. - Verify-trip (marginal comp + HF_DE_VERIFY_MIN_DB=45): fails at frame 649 (32.2dB < 45), peers abort instead of deadlocking, whole render retries via parallel screenshot, RENDER_OK in 42.6s. - Canary suite 7/7 with the flag off (default paths untouched); producer orchestrator tests 99/99; engine suite 909 passed (14 pre-existing main failures, stash-A/B verified); reorder-buffer abort unit tests. Perf note: capture-only parallel speedup measured 1.38x (W2) / 1.52x (W3) over single-worker DE in the spike; end-to-end numbers on this machine are currently noisy (separate-browser init overhead + bench load) — clean benchmarks before any default routing change. The flag stays explicit opt-in; promoting it into the router replaces the #2026 W=1 pin for the same cohort. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
8fee20a525 |
fix(engine): log composition-id attribution on script-failure bail too
Address max-effort code-review finding on PR #2045 (confirmed, not addressed by the earlier review-gap commit): the script_failure bail path skipped the composition-id enumeration entirely, so a render with multiple sub-compositions sharing a failed script only logged the raw failed URL(s), never which composition(s) were still waiting on it — a real observability regression versus the pre-#2045 behavior, which always logged the missing-id list on any non-ready outcome. Now enumerate unregistered composition ids unconditionally and log them alongside whichever reason (script_failure or natural timeout) fired. |
||
|
|
8ba3c33915 |
fix(engine,producer,cli): close review gaps in sub-timeline fail-fast
Address PR #2045 review feedback: - Share a SubTimelineWaitOutcome type (engine) end-to-end instead of widening to string across CapturePerfSummary / RenderPerfSummary / telemetry, so the three layers can't drift. - Dedupe scriptLoadFailures on push — a 4xx response and its trailing requestfailed both recorded the same URL, doubling the failed-URL list in the fail-fast warning. - Thread the sub-timeline-wait outcome into render_error (not just render_complete): a render that fail-fasts and then fails downstream (pollVideosReady, extract, encode) previously dropped this signal on the floor. dedupPerfs is now function-scoped so the catch path can read it, same treatment as the existing captureAttempts array. |
||
|
|
54359f3d6a |
fix(engine): fail-fast the sub-composition timeline wait when a script 404s
pollSubCompositionTimelines waits for every [data-composition-id] host to
register window.__timelines[id]. When the script carrying that registration
fails to load (404 / request failure), the registration can never arrive —
but the poll still burned the full playerReadyTimeout (45s), then warned and
shipped a silently animation-less render. Wild scale: the capture-setup
histogram over 30 days of local renders decays smoothly (402/503/364/282/191
per 5s bucket) then spikes to 705 at the 45s bucket — ~1,000 renders/month
across 402 distinct users, ~15 user-hours of pure waiting.
- Sessions now record failed SCRIPT resources (requestfailed + HTTP>=400
response, listeners that already existed for diagnostics) in
session.scriptLoadFailures.
- pollSubCompositionTimelines takes a failure getter and cuts the wait to a
2s grace once any script failed, with a loud warning naming the URL(s).
Late-registering fetch-async comps are unaffected: no script failure means
the full timeout still applies, and a registration landing inside the
grace window still wins (tested).
- Outcome telemetry: session.subTimelineWaitOutcome ("ready" | "timeout" |
"script_failure") -> CapturePerfSummary -> RenderPerfSummary.subTimelineWait
(worst across sessions) -> render_complete sub_timeline_wait, so the wild
rate becomes directly trackable instead of setup-histogram forensics.
Validation: the discovery comp (0768f038, its animations.js unreachable)
drops from ~72s to 23.1s total — poll cut at 2.1s with the script named;
healthy comp reports "ready". Canary suite 7/7 (PSNRs identical). 4 new
poll unit tests; engine suite 907 passed (14 failures are PRE-EXISTING on
main at v0.7.42 — 18 fail on a clean checkout, stash A/B verified).
tsc/oxlint/oxfmt clean.
Corpus note: 258/1,762 corpus comps (14%) reference local scripts missing
from the corpus fetch — their historical eval INIT timings measured this
timeout, not the engine. Capture-stage ratios remain valid (both paths paid
it equally).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
5f9ee0b678 |
fix(cli,engine): close review gaps in Chrome resolution fix
Address PR #2082 review feedback: - Route studio thumbnail + render call sites through preferManagedChrome so studio renders no longer silently fall back to whatever system Chrome happens to be installed. - `hyperframes browser ensure` now resolves through the same preferManagedChrome path render uses, so it reports what render will actually pick instead of any system Chrome it happens to find. - Point the unsupported-Chrome fallback log at `browser ensure --force` instead of `doctor`, which doesn't check Chrome/drawElement capability. - Fix stale findFromCache comment: the HF pin is now a Dev-channel build that can be newer than a user's puppeteer-cache Stable install. |
||
|
|
8854bad8f9 |
fix(engine,cli): resolve drawElement to a Chrome build that actually has it
canvas.drawElementImage is an unlaunched Dev/Canary-only Blink feature
(~151+). The CLI's pinned CHROME_VERSION fallback was still 131.0.6778.85 —
a puppeteer 24→25.2.1 bump that pinned it to Chrome Dev 151.0.7912.0 was
written on 2026-06-29 but never merged (orphaned local commit, no PR). Any
render on that pin, or on the shared puppeteer-cache binary, or on system
Chrome (Stable, no drawElementImage at all) got a canvas.getContext("2d")
missing the method and crashed mid-capture with "ctx.drawElementImage is
not a function" instead of falling back (HF#2060).
Three changes:
- Bump puppeteer/puppeteer-core to ^25.2.1 across every package that
depends on it, and CHROME_VERSION to 152.0.7928.2 (today's Dev channel;
confirmed via direct probe to implement drawElementImage, unlike 131).
- `ensureBrowser({ preferManagedChrome: true })`, always used by `render`:
resolve straight to our pinned/cached build, skipping both the shared
puppeteer-cache preference and system Chrome. Rendering shouldn't depend
on whatever arbitrary Chrome a machine happens to have — that's exactly
how this regressed (any Mac with Chrome.app installed bypassed the CLI's
pin entirely).
- A runtime capability probe in the engine, right before any other
drawElement work: if `drawElementImage` isn't a function on the injected
canvas, route to the existing screenshot-fallback gate instead of
crashing. This is the real backstop — it protects every resolution path
(env override, stale cache entry, a future Chrome regression), not just
the ones `preferManagedChrome` reaches.
Verified end-to-end: rendering against chrome-headless-shell 131 (confirmed
to lack drawElementImage) now falls back cleanly and produces a valid MP4
instead of crashing; rendering against a capable build still engages
drawElement normally. 922 engine tests + 1373 CLI tests pass.
Fixes #2060.
|
||
|
|
6f8bf5f364 |
fix(engine): resolve relative data-start references for audio tracks (#2062)
parseAudioElements read data-start with a bare parseFloat, so a relative reference (data-start="introClip", the documented 'start when that clip ends' pattern) resolved to NaN. The mixer then silently dropped the track, rendering the whole segment as pure digital silence — even though the SAME reference on the sibling <video> placed the visual correctly (#2030 taught parseVideoElements/parseImageElements to resolve refs; audio never learned). Root fix, single source of truth: extract the Node-side reference resolver out of videoFrameExtractor into referenceResolver.ts and use it in parseAudioElements for both <audio> and <video data-has-audio> tracks. Now every media parser resolves relative timing identically, so audio and video cannot drift again. The two near-identical parse loops share one builder; end stays a numeric read (mixer derives real length downstream), NaN-guarded. Verified end-to-end: a composition with <audio data-start="clipId"> now renders an audio stream that is silent before the referenced clip ends and audible after (matches the numeric-start control); previously the output had no audio stream at all. 78 engine media tests pass (4 new). |
||
|
|
5b9b71df25 |
fix(producer): suppress GSAP call side effects during render seeks (#2037)
* fix(producer): suppress GSAP call side effects during render seeks * fix(core): preserve GSAP root render nudge safely |
||
|
|
4a36655b2b |
fix(engine): resolve relative data-start references in video-frame extraction
* fix(engine): resolve relative data-start references in video-frame extraction <video data-start="intro"> (a relative reference to another clip's end) is resolved by the browser runtime but parseVideoElements/parseImageElements did a raw parseFloat, yielding NaN start/end. The FrameLookupTable active-window checks (start <= t <= end) are then always false, so the clip is never injected and composites BLANK in the final render — while lint/validate/inspect/snapshot and the live preview all look fine. The docs' Relative Timing section teaches exactly this pattern on <video>. Share the pure reference-syntax parser (parseStartExpression) out of the runtime resolver into @hyperframes/core, and resolve references in the extractor against the linkedom document it already holds: a reference resolves to the target clip's resolved start + its duration (data-duration or data-end) + offset, mirroring the runtime. Cycle-guarded; an unknown target or unknown duration falls back to the target's start / 0 (never NaN), matching runtime semantics. Natural-media-duration-only targets aren't known at parse time (same limit as the runtime's fallback). parseImageElements gets the same fix. Runtime resolver behavior is unchanged (its 25-case suite still passes). * chore: re-trigger CI to refresh a stuck CodeQL aggregate check |
||
|
|
76204ec630 |
fix(engine): pre-create __render_frame__ siblings in initializeSession (#2006)
* fix(engine): commit render-frame siblings with a visual BeginFrame at init Chunk-lambda renders drop a periodic near-black frame — one every chunk_frames/worker_count frames (every 60 on a 4-worker single-video chunk), YAVG ~22 against YMAX ~240 in signalstats. Local single-process renders don't show it because they don't run under BeginFrame. It's the isNewImage branch in injectVideoFramesBatch: the first time a session paints a given videoId there's no __render_frame__ sibling yet, so it creates the <img> on the spot (createElement + insertBefore) right before capture. Under HeadlessExperimental.BeginFrame the compositor doesn't have that fresh layer in the immediately-next frame, so the first captured frame per session paints only body background + already-composited overlays. Each lambda worker is its own session, hence the worker-boundary periodicity. Pre-create the hidden sibling at the end of initializeSession, then drive one non-capture visual BeginFrame (noDisplayUpdates: false) to composite the new layers before the first real capture. The warmup ticks are noDisplayUpdates: true (they advance the clock but don't paint) and the per-frame seek doesn't tick, so this explicit visual frame is what actually commits the layers; its tick sits in the gap between warmup and frame 0 so ticks stay monotonic and no render frame is consumed. Every subsequent inject then takes the hasImg=true (src-update) path; the isNewImage branch stays as a fallback for callers that don't go through initializeSession. * fix(engine): place the render-frame commit tick before the liveness probe The commit tick at init sends its BeginFrame at `beginFrameTimeTicks - 1·interval`. The producer's liveness probe then fires right after init at `beginFrameTimeTicks - 5·interval` — an earlier tick. Per-session BeginFrame time has to be monotonic, so the probe running backwards past the commit tick stalls chrome-headless-shell indefinitely; the engine reads that timeout as a SwiftShader heavy-layer stall and routes the render to screenshot capture, which then dies relaunching and hangs the shard to the job timeout. Reproduced on a native x86 SwiftShader host and bisected: with the commit tick present the probe times out even with zero render-frame siblings created, so it's the tick ordering, not layer count. Moving the commit tick to `-6·interval` (below the probe, above the warmup ticks) keeps warmup < commit < probe < capture monotonic and clears the stall on every affected comp — sub-composition-video, chat, style-5-prod — while a healthy comp (style-18-prod) is unchanged. The commit tick itself is untouched, so the black-frame fix it exists for still holds. |
||
|
|
337d0b51bc |
refactor(engine,producer): adopt requestPaint contract, retire autoAlpha rewrite (#2021)
* refactor(engine,producer): adopt requestPaint contract, retire autoAlpha rewrite crbug 529829538 was closed "working as intended": the html-in-canvas API's contract is mutate -> canvas.requestPaint() -> await the canvas paint event -> drawElementImage, which refreshes the subtree's paint records including compositor-applied properties. Verified on the pinned 151 floor and 152 canary: root opacity, root filter, nested group opacity, and child transforms (incl. will-change-promoted) all capture exactly; the root element's own TRANSFORM is the one property still never baked. - Paint invalidation: all three paint-wait sites (serial capture, worker produce, batch produce) now call canvas.requestPaint() when available and fall back to the __hf_de_tick sentinel background toggle on builds without it. The 250ms unsynchronized-draw safety net is unchanged. - Root-opacity ratio correction REMOVED (all three draw sites + base-opacity recording at injection). Since 151 the paint wait bakes current root opacity into the snapshot as pixel alpha, so the ratio correction DOUBLE-APPLIED animated root fades: a root-fade A/B tripped the runtime self-verify at 30.1dB (frame 24, ~0.92 expected vs ~0.85 rendered). Post-removal the same comp self-verifies at inf and matches the screenshot render at PSNR=inf. The root TRANSFORM correction stays — verified still required. - autoAlpha rewrite machinery DELETED: the opt-in opacity->autoAlpha tween rewrite (default-off since the retraction fix; measured ~28dB damage on comps whose fades it touched), its flush-time transparent-target hiding, the __HF_FAST_CAPTURE_AUTOALPHA__ flag plumbing, and the deferral-time retract/re-assert dance. The stub keeps tween-target tracking (3D projection + at-risk scans depend on it). Validation: canary suite 7/7 with PSNRs identical to baseline (58.30 / 43.13 / 54.15 dB); root-fade A/B PSNR=inf vs screenshot; engine suite 905 passed (1 pre-existing color-grading failure); tsc/oxlint/oxfmt clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(engine,producer): review fixes — gate opacity correction by paint mechanism Max code-review findings on the requestPaint adoption: - Root-opacity ratio correction RESTORED, gated per frame on how the paint was produced: it applies on BeginFrame (sync=false) captures and on builds without canvas.requestPaint() — the two paths where the snapshot holds the root's load-time opacity — and is skipped only on requestPaint-driven paints, where the snapshot bakes the current opacity and the ratio double-applies (the proven 30.1dB root-fade failure). Base opacity is recorded at injection again. - Invalidation extracted to a page-scope helper (__hfDeInvalidate, installed by injectDrawElementCanvas) shared by all three paint-wait sites: sentinel toggle ALWAYS (a paint is guaranteed even if requestPaint elides one on a clean subtree) + requestPaint() in a try/catch (a throwing implementation degrades to sentinel-only instead of rejecting the capture). Returns whether requestPaint ran, feeding the opacity-correction gate. Also removes the triplicated inline block and its three anonymous `as T` casts. - HF_FAST_CAPTURE_AUTOALPHA now logs a retirement warning instead of being a silent no-op (the deleted rewrite's comment documented it as an operator escape hatch). - Batch producer docstring updated (still described the tick-toggle-only paint wait); stub tween observer reshaped to a void fn (observeTweenCall) so no arg-rewriting seam survives. Validation: canary suite 7/7 (58.30/43.13/54.15dB, d95f20b6 clean); root-fade A/B self-verify 4x inf + whole-video PSNR=inf; engine suite 905 passed (1 pre-existing); tsc/oxlint/oxfmt clean; stub regenerated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: WaterrrForever <miao.yang@heygen.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b26c27576b |
feat(engine,producer,cli): verify video comps via deferred DE init + capture p50 (#2015)
* feat(engine,producer,cli): verify video comps via deferred DE init + capture p50 Closes the two biggest gaps in the first day of v0.7.38 wild data: 88% of drawElement renders (video comps initialized via probe sessions) ran with self-verification unarmed, and speedup was measurable on only 3 of 76 renders. - Deferred drawElement init: probe sessions initialize before video extraction, so they have no frame injector — ground-truth screenshots would capture black <video> boxes, and verification skipped the whole comp. DE init now stops after the gates for injector-less video comps (deInitDeferred; autoAlpha flag retracted in case no path completes it) and completeDeferredDrawElementInit finishes verification + canvas injection + worker-encode at capture time, once prepareCaptureSessionForReuse has attached the injector. Validated end-to-end: a probe-path video comp now arms 4 ground-truth frames with real video pixels (3x inf + 64.7dB) and renders drawElement verified. - capture_p50_ms: per-frame capture durations are sampled (capturePerf.frameMs; batch frames get the batch mean) and the median ships as CapturePerfSummary.p50TotalMs -> RenderPerfSummary.captureP50Ms -> render_complete capture_p50_ms. Unlike capture_avg_ms it is immune to first-frame warmup and stage-setup amortization — smoke: avg 15ms vs p50 8ms on the same render, p50 matching the measured steady-state floor. Dashboard speedup tiles can drop their frame-count floor once this ships. - video_count on render_complete: segments speedup by video-injection comps (whose per-frame gain is legitimately lower) vs pure-graphics. Canary suite 7/7; engine suite 905 passed (1 pre-existing upstream failure); tsc/oxlint/oxfmt clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(producer): complete deferred drawElement init on the disk capture path Review (miga): a probe-initialized video comp falling back to the disk path kept deInitDeferred and silently stayed in screenshot mode — a regression for PRODUCER_EXPERIMENTAL_FAST_CAPTURE=true renders that previously ran drawElement there. Complete the deferred init on the sequential disk path under the same explicit-opt-in test the orchestrator clamp uses; default-on renders stay on the screenshot baseline (this path has no drain-time self-verification, per the #1998 confinement rule). Validated: video comp + PRODUCER_ENABLE_STREAMING_ENCODE=false + explicit opt-in logs "(deferred drawElement init)" completion on capture_disk and renders correct video pixels. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1005703441 |
feat(engine,producer,cli): drawElement release telemetry on render_complete (#2002)
Default-on drawElement ships with a runtime self-verification net (#1998); this makes its in-the-wild behavior observable. Every render_complete event now answers: which capture mode actually ran, why drawElement disengaged when it did (compile gate / producer clamp / engine init gate), whether the self-verify net fired and why, and how much margin verification had. Follows the static-dedup telemetry pattern: engine session fields → CapturePerfSummary → RenderPerfSummary.drawElement → snake_case props on render_complete. New event props: de_capture_mode, de_compile_gate, de_clamp_reason, de_gate_reason, de_worker_encode, de_verify_armed, de_verify_checked, de_verify_min_db (margin above the 32dB threshold — drift here is the early-warning signal before fallbacks spike), de_verify_init_ms, de_self_verify_fallback, de_fallback_reason, de_blank_suspects, de_blank_deterministic_accepts, de_blank_recaptures, de_boundary_frames, de_ncpr_fallbacks. Validated end-to-end on live renders: drawelement path reports mode/verify counters/minDb/init cost; a blur-gated comp reports mode=screenshot + gate_reason=css_effect:filter; a forced verification failure reports self_verify_fallback=true + fallback_reason=psnr. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ec06f4bf89 |
feat(engine,producer): drawElement fast-capture default-on with runtime self-verification safety net (#1998)
* feat(engine,producer): drawElement fast-capture default-on with runtime self-verification safety net Flip useDrawElement + worker-encode defaults on (HF_DE_BATCH default 4), clamped in resolveConfig to hosts where drawElement can engage (macOS + hardware-GPU browser) so page-side shader compositing is untouched everywhere else; explicit env opt-in keeps attempt-and-gate semantics. Safety net makes default-on safe: the compile/init gates catch predictable incompatibility; this catches the intermittent residue no static analysis can see (stale paints, dropped background images, transient blank frames). - engine: captureDeVerificationFrames — K=4 (HF_DE_VERIFY) ground-truth screenshots at init, after gates + armStaticDedup, BEFORE canvas injection (post-injection screenshots show the canvas bitmap, not the DOM). Runs the video-injection hook per sample; double-captures so rAF-driven text counters settle (a single immediate screenshot captures stale text and false-positives). Skips png, <10 frames, implausible __hf.duration (infinite-repeat GSAP sentinel). - producer: guardFrame on both worker-encode drains — rolling-median blank guard with retry-once at drain (byte-identical retry ⇒ deterministic dark frame, accepted; retry save/restores the static-dedup anchor) + ffmpeg PSNR self-verify vs ground truth (HF_DE_VERIFY_MIN_DB, default 32dB; natural agreement ≥45dB, damage ≤25dB). Breach dumps the frame pair to tmpdir and throws DrawElementVerificationError. - orchestrator: one-shot retry — on verification error the whole render re-runs with forceScreenshot (slower, never wrong); telemetry flag deSelfVerifyFallback. - tooling: de-canary-suite.sh (7-comp release gate with expected verdicts), de-gatecheck.sh (init-only corpus routing classifier), we-render.mjs. Validated: canary suite 7/7; 611-comp routing sample 54% drawelement / 37.5% gated / 8.3% comp-defect; 12/12 risk-band renders clean on bare defaults (48/48 verify samples); engine suite 888 passed; caught two real intermittent damage classes in the wild (background-image drop, root-props offset) that previously shipped silently. Kill switches: PRODUCER_EXPERIMENTAL_FAST_CAPTURE=false, HF_DE_WORKER_ENCODE=false, HF_DE_BATCH=0, HF_DE_VERIFY=0. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(engine,producer): harden the drawElement self-verification net (max code-review findings) 15 confirmed findings from the adversarial review of the default-on flip; the load-bearing five: - Ground-truth capture no longer scrubs GSAP state: seek(0) + forced frame FIRST (lazy .from()/overlap tweens record start values on first seek — mid-timeline scrubs corrupted them for the whole render, and since DE frames and truth shared the corruption, PSNR passed on damaged output), then ascending even-spread fractions, page left at frame 0. - Default-on drawElement is confined to the verified path: resolveConfig requires worker-encode (the drain that runs the net), the orchestrator disengages the default when the render takes the disk path or parallel capture (no drain verification there), and closes a drawElement-initialized probe session rather than letting the unverified path reuse it. Explicit PRODUCER_EXPERIMENTAL_FAST_CAPTURE=true keeps old attempt-and-gate behavior. - Blank-frame retry can no longer splice wrong-frame pixels: recapture goes through recaptureDrawElementFrameForVerify — no static-dedup shortcut (lastEncodeResult runs ahead of the drain) and no "No cached paint record" screenshot fallback (post-injection that captures the canvas = the LAST drawn frame); any recapture failure falls back the whole render. - Verify indices derive from the producer-resolved duration (CaptureOptions.compositionDurationSeconds) instead of raw __hf.duration, so samples always land inside the drained range. - The platform clamp accepts "auto" GPU mode — the stock CLI resolves auto, and the literal-"hardware" clamp made default-on a no-op for the primary audience (masked in validation by explicitly-set env). Also: NaN-safe env parses (HF_DE_VERIFY / HF_DE_VERIFY_MIN_DB / HF_DE_BATCH); video comps skip verification when the session has no frame injector (probe sessions — black-video truth false-positived); psnr infrastructure failures skip the sample instead of failing the render; boundary-saturated sample indices are skipped; shader-transition comps prefer page-side compositing over default drawElement and compile-gated comps get page-side compositing restored; observability.clearFailure un-brands the recovered first streaming attempt; canary suite exempts known-marginal "any" comps from the cross-path PSNR gate; dead we-render options removed; clamp tests pin their env. Validated: canary suite 7/7; auto-GPU bare render engages the full stack; disk-path and worker-encode-off renders disengage default drawElement; malformed HF_DE_VERIFY_MIN_DB still verifies at the default threshold; engine suite 890 passed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(engine,producer): review fixes — PSC intent, verify-threshold clamp, fail-closed canaries Addresses miguel-heygen's review on #1998: - Page-side compositing restore preserves explicit caller intent (blocker): resolveConfig now records pageSideCompositingAutoDisabled only when IT turned page-side compositing off because drawElement was on; the compile-time drawElement gates restore page-side compositing only when that flag is set. An explicit enablePageSideCompositing:false from the programmatic API or HF_PAGE_SIDE_COMPOSITING=false stays off. Pinned by two config tests. - HF_DE_VERIFY_MIN_DB clamped to [10, 60] with a warning on out-of-range values: below ~10dB the check passes severe damage; above ~60dB natural encoder differences force a screenshot fallback on every verified render. - de-canary-suite.sh + de-gatecheck.sh run under set -euo pipefail with explicit `|| true` on expected-nonzero commands (render exits handled by the suite's own checks, grep no-match, kill/pkill/wait races) and a hard FAIL when the PSNR compare produces no value — release canaries fail closed. Full suite re-run green (7/7) under the new flags. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d5ecb013d7 |
perf(engine,producer): batch N drawElement frames per CDP round-trip (HF_DE_BATCH) (#1928)
Amortizes per-frame CDP protocol overhead (~3.5-9ms/frame) by looping seek -> paint-wait -> drawElementImage -> createImageBitmap in ONE page.evaluate for runs of consecutive frames; bitmaps still post to the encode worker per frame. Validated on 19 stratified DE comps: median 1.20x on top of worker-encode (to 1.56x), zero damaged frames, edge comps (static-dedup-heavy, clip-cut) bit-identical; mid-batch failure re-captures via the per-frame path (screenshot-fallback semantics preserved). Off by default; opt in with HF_DE_BATCH=4. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0e58344dca |
feat(engine): frame-capture core — fast-capture routing, worker-encode, dedup extension (#1919)
* feat(engine): drawElementImage capture service * chore(ci): ignore drawElementService exports pending upstack consumers Fallow's per-PR audit diffs against the merge base, so the bottom of the fast-capture stack (#1917) sees drawElementService's exports as unused — their consumers (frameCapture) land in #1919, two PRs upstack. ignoreExports entry documents this and can be dropped once #1919 merges. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(engine): 3D projection + compositor-effect risk gate * fix(engine): gate filter drop-shadow wherever blur gates (review) detectCssEffectRisk documented drop-shadow as a ~29dB damage case but only detected blur( in its three scan paths — a drop-shadow comp stayed on the fast path despite the gate's own correctness contract. Detect drop-shadow( in computed styles, stylesheet rules, and GSAP tween vars, pinned by a focused test that runs the real page-side closure against a DOM shim (computed / stylesheet / tween coverage + blur regression + effect-free null). Addresses miguel-heygen's blocker on #1918. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(engine): frame-capture core — fast-capture routing, worker-encode, dedup extension # Conflicts: # packages/engine/src/services/screenshotService.ts * fix(engine): document HF_FORCE_DRAWELEMENT as diagnostic-only; make armStaticDedup idempotent (review) Addresses miguel-heygen's blockers on #1919: - HF_FORCE_DRAWELEMENT promoted from a stale "SCRATCH/Uncommitted" comment to a documented diagnostic flag: it exists for upstream-Chromium repro work (gate-vs-API isolation, crbug 521861819 149-vs-151) and R&D on gated effect classes; renders under it may be damaged BY DESIGN since it bypasses gates whose thresholds encode measured damage. Never production; the safety-net blank guard also stands down under it so diagnostic frames arrive unmodified. - armStaticDedup is now idempotent: the drawElement init path arms dedup before canvas injection, then initializeSession called it again — the second run overwrote the armed state with skipReason="capture_mode" (captureMode is "drawelement" by then), producing contradictory telemetry (armed frames + a skip reason), and re-ran the verification seeks on the fallback path. It now no-ops once staticFrames or a skip decision exists. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
4749fe5716 |
feat(engine): 3D projection + compositor-effect risk gate (#1918)
* feat(engine): drawElementImage capture service * chore(ci): ignore drawElementService exports pending upstack consumers Fallow's per-PR audit diffs against the merge base, so the bottom of the fast-capture stack (#1917) sees drawElementService's exports as unused — their consumers (frameCapture) land in #1919, two PRs upstack. ignoreExports entry documents this and can be dropped once #1919 merges. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(engine): 3D projection + compositor-effect risk gate * fix(engine): gate filter drop-shadow wherever blur gates (review) detectCssEffectRisk documented drop-shadow as a ~29dB damage case but only detected blur( in its three scan paths — a drop-shadow comp stayed on the fast path despite the gate's own correctness contract. Detect drop-shadow( in computed styles, stylesheet rules, and GSAP tween vars, pinned by a focused test that runs the real page-side closure against a DOM shim (computed / stylesheet / tween coverage + blur regression + effect-free null). Addresses miguel-heygen's blocker on #1918. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
795934adba |
feat(engine): drawElementImage capture service (#1917)
* feat(engine): drawElementImage capture service * chore(ci): ignore drawElementService exports pending upstack consumers Fallow's per-PR audit diffs against the merge base, so the bottom of the fast-capture stack (#1917) sees drawElementService's exports as unused — their consumers (frameCapture) land in #1919, two PRs upstack. ignoreExports entry documents this and can be dropped once #1919 merges. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
644ae384a5 | fix(runtime): satisfy color grading audit | ||
|
|
870964b0cf | feat(runtime): render media color grading shaders | ||
|
|
b7dcb9e2a3 | fix(engine): harden ffmpeg binary resolution | ||
|
|
56d4a7032b | fix(engine): preserve DOM mask visibility state (#1953) | ||
|
|
bb066077b4 |
fix(producer): avoid reviving hidden DOM in HDR layers (#1935)
* fix(producer): avoid reviving hidden DOM in HDR layers * fix(producer): filter transition HDR DOM masks * fix(producer): keep hidden timed descendants masked |
||
|
|
8a3227f548 |
fix(engine): write the audio mix filter graph to a file, not the command line (#1890)
* fix(engine): write the audio mix filter graph to a file, not the command line mixAudioTracks built the ffmpeg -filter_complex argument as one inline string scaling linearly with track count. Reported in the wild at 146 timed audio clips: the resulting command line exceeded the OS length limit and spawn failed with ENAMETOOLONG, dropping audio entirely until the user manually consolidated clips to reduce the count. FFmpeg supports -filter_complex_script specifically for this - the same filter graph read from a file instead of inlined as an argument. The -i pairs for each track still scale with count but stay short and fixed-size each, so the one component that actually grew unbounded (the filter string) no longer sits on the command line at all. The temp file is cleaned up immediately after ffmpeg exits, matching the existing sibling temp-file convention in audioVolumeEnvelope.ts. Verified end-to-end against a real ffmpeg binary (not just mocked): a two-track mix produced correct output audio with no leftover temp files. * fix(engine): create audio filter scripts safely |
||
|
|
1a7002f208 |
perf(engine): superset extraction for overlapping trims of one source (#1885)
* perf(engine): superset extraction for overlapping trims of one source Cache-missing trims of the same source that are frame-aligned and overlapping decode their union window in ONE ffmpeg pass; each trim's frames are materialized by hardlinking the superset frames with renumbered names (copy fallback on EXDEV). Byte-identical to per-trim extraction on CFR sources (verified by content hash in the A/B run), ~2x less decode+encode work for typical overlapping trims, and sparse-keyframe sources pay the keyframe seek once instead of once per trim. Disjoint or misaligned trims keep the direct path; any union failure falls back to per-trim extraction. Also: warm renders (zero cache misses) skip the extraction-cache GC sweep instead of paying a full cache size scan. * fix(engine): superset review hardening - clustering, abort, cache-fs temp, gc staleness - Partition each source's trims into overlap-connected components before the union check, so one disjoint outlier no longer collapses the whole bucket to direct extraction (pinned by a 3-of-4-overlap test). - On abort, the superset fallback no longer re-runs every member through direct extraction (N doomed ffmpeg spawns); the cancellation surfaces per member instead. - The superset temp dir moves onto the cache filesystem when the cache is active so member hardlinks into partial dirs cannot EXDEV-copy and silently multiply disk usage; its .partial- name puts crashed leftovers under the GC's aged-partial sweep. - GC staleness fallback: a .hf-last-gc marker is stamped per sweep and all-hit renders sweep anyway once it is older than 24h, so 100%-warm workloads still reclaim space (pinned by a stale-marker test). |
||
|
|
48f158a0c2 |
perf(engine): one-pass SDR-to-HDR extraction with cache-key transform (#1902)
* perf(engine): one-pass SDR-to-HDR extraction with cache-key transform Mixed-HDR compositions converted each SDR source with a full libx264 re-encode (convertSdrToHdr) before extraction. The BT.709 to BT.2020 colorspace remap now runs as a filter inside the extraction pass itself; convertSdrToHdr and the _hdr_normalized intermediate are deleted. Same shape as the earlier one-pass VFR change. Also fixes a cache-poisoning bug this exposed: the HDR preflight rewrote entry.videoPath AFTER the cache-key snapshot, so a mixed-HDR render cached converted frames under the plain source key and a later SDR render of the same trim would have served HDR-tinted frames. The cache key now carries an optional transform discriminator; keys without a transform stay byte-compatible with existing entries. * fix(engine): attribute SDR-to-HDR extract failures, pin filter-order intent Review hardening for one-pass SDR-to-HDR: - ffmpeg failures now carry an 'SDR→HDR conversion failed (colorspace filter in extract pass)' prefix when the remap is in the chain, so a filter-less ffmpeg build fails loudly with attribution instead of a generic extract error. - Comments pin the fps-before-colorspace ordering intent and mark sdrToHdrTransfers as the canonical read for both the cache key and extraction options. - Cross-render cache-poisoning regression test now compares frame BYTES across the cache boundary: mixed-HDR render then plain-SDR render of the same trim must produce different pixels, and a repeat plain render must hit the plain entry with byte-identical frames. |
||
|
|
34590649a0 |
perf(engine): extraction cache on by default with atomic publish and LRU gc (#1901)
* perf(engine): extraction cache on by default with atomic publish and LRU gc Warm re-renders now skip source-video frame extraction entirely (video_extract 400ms -> 13ms on a 4-video composition; outputs are pixel-identical, PSNR inf). What made default-on safe: - Atomic entry publish: frames extract into a unique .partial-<pid>-<uuid> dir, the completion sentinel is written there, and the dir is renamed into the final key atomically. Concurrent renders sharing a cache can duplicate work but can never serve a torn entry (previously documented as single-writer only). - Size-capped LRU gc: best-effort sweep after extraction evicts oldest-used entries past a 2 GiB default budget (HYPERFRAMES_EXTRACT_CACHE_MAX_MB) and clears crashed writers' partials. Entries younger than 60 min are never evicted so live renders keep their frames. - Default cache dir: <tmpdir>/hyperframes-extract-cache-<uid>. Opt out with HYPERFRAMES_EXTRACT_CACHE_DIR=off (or none/false/0); a non-writable dir degrades to uncached with a single warning instead of failing the render. * fix(engine): harden extraction cache publish and surface cache ops signals Review hardening for the default-on extraction cache: - Bypass the cache for HDR-converted intermediates: the key snapshot describes the original source, so publishing converted frames under it would poison later plain-SDR renders of the same trim. (The follow-up transform-keyed change re-enables caching for these.) - publishCacheEntry TOCTOU: adopt a concurrent writer's completed entry both before removing an apparently-stale dir and after a failed retry rename, so a winner's publish is never destroyed or reported as a failure. - Observability for the failure paths: cachePublishFailures, cacheGcEvictions, cacheGcBytesFreed, and cacheAgedPartialsCleared on ExtractionPhaseBreakdown; gcExtractionCache now returns sweep stats. * fix(engine): sweep superseded cache generations in gc After a SCHEMA_PREFIX bump, old-generation entries (hfcache-v2-*) no longer matched the sweep's prefix filter and would orphan their disk forever. The gc now matches any hfcache-v* generation; superseded entries never receive sentinel touches, so the LRU evicts them first. |
||
|
|
8d64d48e4a |
perf(engine): dedupe identical extractions within one render (#1900)
* perf(engine): write PNG frames at compression_level 1 Extracted video frames are render-scoped temp files read once during capture, so zlib effort above level 1 buys nothing. Measured 3.3x faster on 60s of 1080p H.264 to PNG (11.4s to 3.5s) and 5.4x on a 20s vp9-alpha webm (4.3s to 0.79s), for ~14% larger temp files. * perf(engine): one-pass VFR extraction with -fps_mode cfr VFR sources (screen recordings, phone videos) were re-encoded to CFR with libx264 and then extracted in a second ffmpeg pass. Extraction now runs a single pass with -fps_mode cfr -r <fps>. Same frame counts on the VFR regression fixtures (120/120 mid-seek, 297-303 full file), one less x264 generation of quality loss, ~3.4x faster on VFR inputs. convertVfrToCfr and the _vfr_normalized intermediate are deleted. The full-VFR test's byte-identical duplicate-frame cap is retired with cause: the fixture has no source frames for 40% of its timeline, so held frames are correct; the two-pass path only scored under it because x264 encoder noise made frozen frames hash differently. The freeze regression (missing frames) stays pinned by the frame-count windows. * docs(engine): pin vfrPreflightMs definition change after one-pass VFR vfrPreflightMs used to time a per-source VFR-to-CFR re-encode; it now times only the cached classification probe and collapses to ~0. Call that out on ExtractionPhaseBreakdown so dashboards keyed on the old threshold semantics migrate to vfrPreflightCount / extractMs. * fix(engine): bump extraction cache schema to v3 for one-pass VFR frames One-pass VFR extraction changes frame CONTENTS for VFR sources while the cache key tuple (path, mtime, size, trim, fps, format) is unchanged, so warm v2 entries holding two-pass frames would keep being served across the deploy boundary. Bumping the schema prefix makes v2 entries inert; affected sources re-extract once. * perf(engine): dedupe identical extractions within one render N <video> elements sharing (resolved path, mediaStart, duration, fps, format) extracted N times; they now share one extraction via an in-flight promise map keyed on that tuple. Duplicate elements receive the shared frame set under their own videoId. This also removes a race where two identical clips on a cache miss wrote the same extraction-cache entry dir concurrently. 3x duplicated 60s 1080p video: 4426ms to 1521ms in the A/B benchmark, one frame set on disk. * fix(engine): attribute shared-extraction failures to the dedupe leader When a deduped extraction fails, every follower reported the leader's error verbatim under its own videoId, reading as N independent failures in traces. Follower errors now carry a '[shared extraction, leader <id>]' prefix so the fan-out is traceable to one root failure. |
||
|
|
7860583341 |
perf(engine): one-pass VFR extraction with -fps_mode cfr (#1899)
* perf(engine): write PNG frames at compression_level 1 Extracted video frames are render-scoped temp files read once during capture, so zlib effort above level 1 buys nothing. Measured 3.3x faster on 60s of 1080p H.264 to PNG (11.4s to 3.5s) and 5.4x on a 20s vp9-alpha webm (4.3s to 0.79s), for ~14% larger temp files. * perf(engine): one-pass VFR extraction with -fps_mode cfr VFR sources (screen recordings, phone videos) were re-encoded to CFR with libx264 and then extracted in a second ffmpeg pass. Extraction now runs a single pass with -fps_mode cfr -r <fps>. Same frame counts on the VFR regression fixtures (120/120 mid-seek, 297-303 full file), one less x264 generation of quality loss, ~3.4x faster on VFR inputs. convertVfrToCfr and the _vfr_normalized intermediate are deleted. The full-VFR test's byte-identical duplicate-frame cap is retired with cause: the fixture has no source frames for 40% of its timeline, so held frames are correct; the two-pass path only scored under it because x264 encoder noise made frozen frames hash differently. The freeze regression (missing frames) stays pinned by the frame-count windows. * docs(engine): pin vfrPreflightMs definition change after one-pass VFR vfrPreflightMs used to time a per-source VFR-to-CFR re-encode; it now times only the cached classification probe and collapses to ~0. Call that out on ExtractionPhaseBreakdown so dashboards keyed on the old threshold semantics migrate to vfrPreflightCount / extractMs. * fix(engine): bump extraction cache schema to v3 for one-pass VFR frames One-pass VFR extraction changes frame CONTENTS for VFR sources while the cache key tuple (path, mtime, size, trim, fps, format) is unchanged, so warm v2 entries holding two-pass frames would keep being served across the deploy boundary. Bumping the schema prefix makes v2 entries inert; affected sources re-extract once. |
||
|
|
557a270271 |
perf(engine): write PNG frames at compression_level 1 (#1898)
Extracted video frames are render-scoped temp files read once during capture, so zlib effort above level 1 buys nothing. Measured 3.3x faster on 60s of 1080p H.264 to PNG (11.4s to 3.5s) and 5.4x on a 20s vp9-alpha webm (4.3s to 0.79s), for ~14% larger temp files. |
||
|
|
df221c1fd6 |
fix(engine): scale static-dedup verification density with run length (#1903)
* fix(engine): scale static-dedup verification density with run length Reported symptom: a 10-scene template composition (shared card layout, per-scene text/progress-bar content) rendered scene 1 correctly, but every scene after that had its text/progress-bar card missing from the final MP4 -- even though snapshot and validate showed correct per-scene content when seeking directly to those timestamps. Setting HF_STATIC_DEDUP=false fixed every scene. Render log showed a large, mostly-reusable static-frame run engaging (2430 frames, 34% reusable). verifyStaticFramesSafe already does a real, pixel-exact comparison (anchor vs. candidate screenshot) before trusting a predicted-static run -- the reuse mechanism itself is correct and already regression- locked (frameCapture-staticDedupIndex.test.ts). The gap was sample density: interior checks per run were capped at a flat min(sampleCount, 8) points, so the stride between checks grew with the run's span. A 2000+ frame run (plausible for a 10-scene comp where computeStaticFrameSet's GSAP-tween-only interval walk can't see whatever mechanism swaps each scene's text) could space checks ~285 frames apart, letting a real content change hide between two verified points and get the whole run wrongly trusted as static. Fix: extract the point-selection into a pure, exported computeStaticVerificationPoints(a, b, sampleCount), and bound the STRIDE by sampleCount (HF_STATIC_DEDUP_SAMPLES) instead of just the point count, so density scales with run length. Short/typical runs are unaffected (the two formulas agree there); long runs get proportionally denser checks. The existing hardCap safety valve is untouched -- if this makes verification too expensive for a pathological composition, dedup still disarms entirely rather than trusting a sparsely-checked set. Test: new frameCapture-staticDedupVerifyDensity.test.ts asserts the max gap between consecutive verification points never exceeds sampleCount on long runs (would fail pre-fix at span=2000/10000), matches the prior stride on short runs, and always includes both run endpoints. Full engine suite (845 tests) passes. * fix(engine): decouple verification density scaling from sampleCount polarity Addresses review feedback on the static-dedup density fix (PR #1903): 1. The prior revision bounded the interior-check STRIDE by sampleCount directly, which inverted HF_STATIC_DEDUP_SAMPLES' polarity: raising it widened the allowed gap between checks instead of narrowing it, and the "raise HF_STATIC_DEDUP_SAMPLES to verify more" log guidance became backwards for exactly the long runs it's meant to help. Fix: introduce a fixed STATIC_VERIFY_REFERENCE_STRIDE (24 frames, independent of sampleCount) that drives the length-scaling behavior -- this alone fixes the original bug (long runs going nearly unverified) regardless of how sampleCount is configured. sampleCount is now purely a per-run point-count FLOOR: raising it only ever increases density, restoring correct, monotonic polarity. 2. hardCap wasn't re-tuned for the new cost model. The old flat 8-point cap cost ~8 checks/run; the new density costs ~span/24 checks/run -- ~103 for the reported 2430-frame run, ~417 for a 10k-frame run. Sizing the budget only off sampleCount (which no longer drives density for long runs) would make a genuinely-static long composition spuriously disarm under the new, more thorough checking. hardCap now also scales with the total predicted-static frame count, with a 3x margin over the expected minimum verification cost. Softened the budget-exhausted log message accordingly -- it no longer prescribes raising sampleCount, which would often just add cost without proportionally raising the now length-driven budget. 3. The 5 existing tests only asserted sample-point geometry (gaps, endpoints, stride shape), not the actual point of the fix -- that a real content change hiding between the OLD sample gaps now gets caught. Added a behavior-level test: mocks pageScreenshotCapture to simulate a transient content change at a frame the pre-fix formula would have skipped (reconstructed locally in the test, commented as historical-only) but the new formula samples, and asserts the real verifyStaticFramesSafe (now exported) detects it via the real computeStaticVerificationPoints -- not a reimplementation. Also added a direct polarity regression test (raising sampleCount past the length-scaled floor must strictly tighten the gap) and reworded the short-run test to reflect the corrected formula. Full engine suite (847 tests) passes. |
||
|
|
eef4690752 |
fix(engine): name the fix in the ffmpeg encode-timeout error message (#1858)
Two independent post-release feedback reports of hitting
ffmpegEncodeTimeout (600000ms default) on long or high-frame-count
renders, both resolved by setting FFMPEG_ENCODE_TIMEOUT_MS to a higher
value and/or PRODUCER_ENABLE_CHUNKED_ENCODE=true — env vars that already
exist and already solve this, but that neither user found from the error
message itself.
appendEncodeTimeoutMessage only stated what happened ("FFmpeg killed after
exceeding ffmpegEncodeTimeout"), not what to do about it. Name both
existing knobs in the message so the fix is immediately visible at the
point of failure instead of requiring a source dive.
One function, six call sites, all fixed at once. Existing tests assert
with toContain, so the appended text doesn't break them; added two
assertions confirming both env var names appear in the message.
|
||
|
|
145c71e837 | fix(engine): parallelize forced screenshot workers (#1848) | ||
|
|
c0c3abf0f1 |
fix(producer): harden capture against timeouts, transient tab deaths, and OOM (#1842)
Four independent capture-infra hardening changes for the P2-5 failure bucket (~15K err / ~7K users): - protocolTimeout auto-scales by device-scaled output area (applied before probe launch, since it's immutable post ppt.launch()). - Single bounded transient retry (MAX_TRANSIENT_CAPTURE_RETRIES=1) on Target closed / Page crashed in the parallel disk-capture path; abort short-circuits before retry. - Narrow OOM classification (Set maximum size exceeded etc., disjoint from transient) → actionable guidance naming output dims. - StreamingEncoder.getExitError() threads FFmpeg's real exit reason into frame-0 encoder-death errors. Render-reliability workstream P2-5. Success measured on PostHog dashboard 1783183. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
24edb15095 |
fix(runtime): auto-infer composition duration for CSS/WAAPI/Lottie so data-duration is optional (#1830)
* fix(runtime): auto-infer composition duration for CSS/WAAPI/Lottie so data-duration is optional The #2 render failure bucket ("Composition has zero duration") accounts for ~27K errors / ~7K affected users over 30 days (PostHog project 356858). Root cause: only GSAP timelines got their duration auto-detected — CSS, WAAPI, and Lottie compositions had no source of truth for total duration unless the author remembered to set data-duration on the root element, and the render engine hard-failed capture when neither was present. Adds getInferredDurationSeconds() to the CSS, WAAPI, and Lottie runtime adapters (packages/core/src/runtime/adapters/*.ts) — each reports the longest finite end time it can discover from its own animations (CSS: computed timing offset by data-start; WAAPI: effect.getComputedTiming().endTime; Lottie: totalFrames/frameRate or the player's own duration). Infinite/ unbounded animations correctly return null and still require data-duration. Wires this into the runtime's existing duration-floor resolution (resolveAdapterDurationFloorSeconds in runtime/init.ts), alongside the existing media-duration and authored-composition floors, so window.__hf.duration becomes positive without any author action for finite-duration non-GSAP compositions. Three.js is unchanged — no AnimationClip/AnimationMixer inspection exists in that adapter, so data-duration remains required there. Tightens frameCapture.ts's zero-duration fast-fail gate to also check hf.duration directly (not just the two authored signals), so a composition mid-inference isn't fast-failed before its adapter-derived duration lands. Adds a new lint rule (root_composition_missing_duration_source) that errors only on genuinely non-inferable cases: no animation signal at all, Three.js without data-duration, or an infinite/unbounded CSS or WAAPI animation without data-duration. Deliberately silent on finite CSS/WAAPI/Lottie animations, since the runtime now infers those — an autofix that "inserts the inferred value" was considered and rejected: every case the rule flags has no derivable value (an infinite spinner has no finite end time; a duration-less Three.js scene has nothing to measure), so any autofix would have to fabricate a placeholder, trading a loud correct failure for a silent wrong-length render. Updates the CSS/WAAPI/Lottie/Three adapter skill docs and the hyperframes-core determinism-rules/data-attributes references to document the new optionality and the runtime mechanism backing it. Verified end-to-end against the real render pipeline (not just unit tests): a CSS-only composition with a finite 3s animation, no GSAP timeline, and no data-duration now renders a correct 3.000s MP4 via `hyperframes render` (previously: "Composition has zero duration" failure). The infinite-CSS negative control still fails fast with a clear diagnostic, matching the new lint rule. Adds a file-level fallow health exemption for lottie.ts's pre-existing `seek` handler — unrelated to this change, but its line numbers shifted when new functions were added earlier in the file, tripping fallow's inherited-finding fingerprint (documented pattern already used elsewhere in .fallowrc.jsonc for the same reason). Known limitation: the static WAAPI usage detector in the lint rule (/\.animate\(\s*[\[$A-Za-z_]/) can miss unusual call shapes; it only affects whether the "no signal at all" branch fires, and errs toward NOT flagging (reducing false positives) rather than over-flagging. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(lint): close 3 correctness gaps in root_composition_missing_duration_source - Strip JS/CSS comments before scanning for GSAP/WAAPI/Three/Lottie/CSS animation signals, so a commented-out `.animate()` call or a commented `animation: ... infinite` rule can no longer satisfy the "has a duration source" check and mask a real zero-duration render failure. - Broaden the WAAPI detection regex to also match the object-literal (PropertyIndexedKeyframes) form of `.animate()`, e.g. `el.animate({ opacity: [0,1] }, { duration: 2000 })`, which the previous character class silently missed. Corrected the adjacent comment that incorrectly claimed this shape "can't be a false negative". - Fix hasInfiniteCssAnimation to stop false-positiving on animation NAMEs that merely contain the substring "infinite" (e.g. `infinite-spin`) by anchoring the `infinite` keyword with hyphen-aware boundaries instead of a bare `\b`. Also makes the longhand `animation-name` + separately declared `animation-iteration-count: infinite` pattern detected consistently. Adds targeted unit tests for each fixed false-positive/false-negative. * fix(runtime): keep finite duration signal when an unbounded animation coexists getInferredDurationSeconds in the CSS and WAAPI adapters returned null outright whenever any animation on the composition was unbounded (infinite iteration count), even when other finite animations on the same composition could still supply a valid duration. This disagreed with the new root_composition_missing_duration_source lint rule, which treats any animation-name as sufficient — so a composition mixing a finite fadeIn with a decorative infinite spin passed lint but still failed at render with "zero duration". Unbounded animations are now skipped when computing the max end time instead of short-circuiting the whole calculation. null is only returned when every animation on the composition is unbounded, i.e. there is no finite signal to fall back on at all. Co-Authored-By: Claude <noreply@anthropic.com> * docs(skills): fix table separator width in data-attributes.md oxfmt flagged the merged Composition Root table from the post-rebase merge of the auto-infer-duration docs onto main's reformatted table — the separator row was one dash short of the header width. * fix(lint): keep infinite-CSS duration rule strict but make its message honest Post-review (Vance): after the finite+infinite adapter fix, the runtime infers a length for a mixed finite+infinite CSS composition, but this lint rule still (intentionally) errors on it — an unbounded animation makes the intended total length ambiguous, so we require explicit data-duration. Keep that strictness (lint is advisory by default; it only blocks under --strict, and data-duration is the one duration signal guaranteed correct across every adapter, known and future). But the message wrongly claimed the render "will fail" — false for the mixed case, where the runtime falls back to the finite animation. Rewrite it to describe the ambiguity honestly, correct the rule's block comment, and add a mixed finite+infinite test asserting it still errors with an honest message. --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
b33d54f54b |
fix(engine): retry probe on pollHfReady zero-duration timeout (#1824)
Renders were failing outright with "[FrameCapture] Composition has zero duration. Runtime ready: false, ..." whenever window.__renderReady didn't flip true within playerReadyTimeout (45s) — most often under host contention (e.g. several renders running concurrently), never from a defect in the composition itself. Confirmed by re-running an affected composition standalone: it succeeded immediately (initMs ~3.5-4.4s vs. the 45s timeout it hit under concurrent load). The probe stage already retries once with a fresh browser session for exactly this class of "succeeds on retry" infra flakiness (frame detachment, disconnects, navigation timeouts, launch failures), but isTransientBrowserError didn't recognize this message, so it fell through to an immediate, unretried failure. Match "Composition has zero duration ... Runtime ready: false" as transient. Left the "Runtime ready: true" case (pollHfReady's fast-fail: no GSAP timeline and no data-duration) unmatched — that's a genuine authoring bug, not a timing fluke, and should keep failing fast. |