mirror of
https://github.com/heygen-com/hyperframes.git
synced 2026-09-03 04:38:33 +00:00
## Summary PR 4 of 5 in the hf#732 decomposition stack. **This is where the bulk of the shader-transition speedup lives** (`~2×` verified — see Empirical validation below). Spreads per-frame DOM capture work across N DOM worker sessions and offloads the per-pixel shader-blend onto a `worker_threads` pool (the pool added in #758). ### Gating The hybrid path is gated by `shouldUseHybridLayeredPath`: - SDR content only — HDR raw-frame sources are fd-bound to one worker (per-worker `dup(fd)` is out of scope here). - `workerCount >= 2`. - Not every frame inside a transition window. When the gate trips, the hybrid loop spawns `workerCount - 1` extra DOM sessions, allocates per-worker scratch buffers, and partitions the frame range into contiguous slices via `distributeLayeredHybridFrameRanges`. Each worker walks its slice; transitions dispatch through the shader-blend pool (with inline fallback). A frame-reorder buffer fences the encoder. Pool teardown is guaranteed via try/finally on both the success and error paths. ### Structural change (heads-up to reviewers) `captureHdrStage.ts` on main was already 921 lines (over the project's 500-line ceiling). Adding the hybrid path on top would push it past 1100 and the local pre-commit hook refuses to stage files past 500. **PR 4 splits `captureHdrStage.ts` into 5 files**: - `captureHdrStage.ts` (orchestrator + cleanup invariants, 469 lines) - `captureHdrResources.ts` (HDR video extraction + image decode + dim probing) - `captureHdrFrameShared.ts` (gating predicates, partitioning, per-scene capture) - `captureHdrSequentialLoop.ts` (legacy single-session loop) - `captureHdrHybridLoop.ts` (new multi-worker path) No behavior change in any pre-existing code path: the sequential loop is byte-equivalent to the previous inline implementation (both consume `captureSceneIntoBuffer` from the shared module, so behavior parity is enforced structurally rather than by comment-keeping). `renderOrchestrator.ts` is intentionally unchanged — the stage computes its own worker budget via `calculateOptimalWorkers` rather than receiving it through the call signature. ## Stack Stacked on top of #758 (PR 3: shaderTransition pool). ## Test plan - [x] 14 new vitest tests in `captureHdrFrameShared.test.ts` pinning the hybrid gating predicate and the contiguous-chunking partitioner — all pass - [x] Producer typecheck clean - [x] oxlint clean ### Empirical validation Mark Witt fixture (Mac, Apple Silicon, hardware GPU, no beginframe): - Published CLI (pre-stack): 2m 12.2s - Cascade CLI (this stack): 1m 07.7s - **Measured speedup: 1.95× on Mac.** (Earlier "2.22×" wording was a projection from per-component micro-benchmarks; the empirical end-to-end number is 1.95× on the validated fixture.) Linux CI confirmation pending top-of-stack regression run. — Vai