Files
hyperframes/packages/producer/tests
Miguel Ángel 2be8a62c00 fix(engine): stop compositing phantom duplicates on captureBeyondViewport (#2607)
* fix(core): stop the async media-metadata rebind once render capture starts seeking

scheduleMetadataDurationHydration re-resolves and can swap the captured
GSAP timeline off a debounced loadedmetadata/durationchange event, fully
uncoordinated with the producer's own per-frame renderSeek calls. When a
full-length <video>'s metadata resolves after capture has already begun
(slow I/O, Docker), this races the deterministic BeginFrame capture loop
and can reflow sub-composition state mid-render, producing phase-offset
duplicate content in captured frames (#2550).

Render-mode duration correction already happens deterministically during
the probe stage before capture starts, so once renderSeek has been called
once there is nothing left for this self-correction to do — gate it off
for the rest of the session.

* fix(core): scope the metadata-rebind guard to actual render/export pages

renderSeek isn't capture-exclusive — Studio's own preview iframe falls
back to it for compositions whose timeline overhangs every native
adapter's duration. Gating the HF#2550 fix on renderCaptureSeekStarted
alone silently disabled the metadata-driven duration self-correction for
that live-scrub case too, where it's still needed. Require the render/
export page signal (window.__HF_EXPORT_RENDER_SEEK_CONFIG, set only by
the producer's fileServer.ts) alongside it, and add a regression test
covering the Studio-preview case.

* fix(engine): stop requesting beyond-viewport capture for video comps that don't need it

Root-caused HF#2550 by reproducing the reporter's public repro end-to-end
(not just the timeline-rebind mechanism from the earlier commits in this
branch) on native Linux: instrumented the actual DOM state during a real
capture session and confirmed the sub-composition never double-mounts —
getBoundingClientRect and the timeline's own local time both match the
single, correct DOM tree throughout. The phantom second copy only exists
in the captured screenshot pixels.

Bisected it to captureBeyondViewport: resolveVideoCaptureBeyondViewport
(#1094's tall-portrait fix) forces `Page.captureScreenshot`'s beyond-viewport
path on for any render with a native <video>, regardless of whether the
page's content actually overflows the declared capture height. On
SwiftShader that beyond-viewport path can composite a stale, vertically
offset paint of the page alongside the fresh one for content that fits
entirely within the viewport — producing exactly the reported phase-offset
duplicate. Disabling captureBeyondViewport (repro's video still present)
eliminates the duplicate outright; re-enabling it reproduces the duplicate
byte-for-byte, isolating it as the actual cause.

Adds pageContentExceedsCaptureHeight, a ground-truth measurement of the
page's actual scrollHeight against the requested capture height, and wires
it into initializeSession to downgrade captureBeyondViewport back to false
once the page is settled and it's confirmed unnecessary — the "reliable
clip predictor" the original #1094 fix's ponytail comment flagged as
missing. This keeps #1094's fix intact for content that genuinely
overflows while closing the SwiftShader ghosting hazard for the (common)
case of video that fits inside its own viewport.

* test(producer): add HF#2550 video+sub-composition regression fixture

Checks in the reporter's confirmed real-world reproduction (media
regenerated via ffmpeg testsrc2, matching their public repro repo) as a
regression fixture, with a golden baseline rendered against the fix.

Verified end-to-end via the project's own Docker regression harness:
- Rendering this fixture with the fix produces the golden baseline
  (clean, single flowchart instance, captureBeyondViewport correctly
  downgraded).
- Direct CLI renders (not through this harness) against unpatched code
  reproduce the reported phantom-duplicate artifact reliably (10/10).

Caveat documented in meta.json: the underlying bug is timing-dependent.
Two harness runs against unpatched code, using this same fixture, did
not reproduce the artifact (0/2) — the harness's in-process render path
apparently doesn't hit the same race window a direct CLI process does on
this host. This fixture is a best-effort regression guard and a
preserved real-world repro, not the sole protection — the deterministic
guard is packages/engine/src/services/screenshotService.test.ts's
pageContentExceedsCaptureHeight unit tests, which exercise the actual
fix logic directly.

Also adds an .gitattributes LFS rule for this fixture's source
index.html (744 KB — carries the real project's embedded base64
assets, over the largefiles hook's 500 KB non-LFS limit).

* fix: route HF#2550 fixture binaries through LFS (were committed raw)

filter.lfs.clean/smudge were locally configured as a no-op "cat" in
this repo's shared .git/config, silently disabling LFS filtering for
every worktree. The previous commit's large binaries (output.mp4,
compiled.html, source index.html, source video) landed as raw blobs
instead of LFS pointers as a result. Ran `git lfs install --local
--force` to restore the correct filter commands, then re-staged the
affected files so they commit as proper LFS pointers.

* fix(engine): address capture viewport review feedback
2026-07-17 11:59:45 -04:00
..
2026-04-27 18:16:09 -04:00

Producer regression test fixtures

Each subdirectory under this folder is a regression fixture for the HTML-to-video pipeline. The harness at packages/producer/src/regression-harness.ts walks every subdirectory, runs the composition, and PSNR-compares the rendered output against a checked-in golden baseline.

Fixture layout

<fixture-name>/
├── meta.json           # name, tags, PSNR threshold, renderConfig
├── src/
│   ├── index.html      # composition entry point
│   └── assets/...      # any locally-referenced media
└── output/
    ├── compiled.html   # golden compiled HTML (validated as a snapshot)
    └── output.mp4      # golden rendered video

meta.json is validated by validateMetadata in src/regression-harness.ts. The required fields are:

  • name (string), description (string), tags (string[])
  • minPsnr (number, dB)
  • maxFrameFailures (integer)
  • minAudioCorrelation (0..1), maxAudioLagWindows (integer ≥1)
  • renderConfig.fps (integer like 30 or a rational string like "30000/1001")

Optional renderConfig fields:

  • format"mp4" (default) or "webm"
  • workers — integer ≥ 1
  • hdr — boolean (default false)
  • variables — JSON object of render-time variable overrides
  • chunkSize — integer ≥ 1 (used by --mode=distributed-simulated)
  • maxParallelChunks — integer ≥ 1 (used by --mode=distributed-simulated)

Generating / updating a baseline

Always inside Docker. Host Chrome / FFmpeg versions drift across distros, so a baseline captured on the host won't match the bytes CI renders.

# From the repo root.
docker build -t hyperframes-producer:test -f Dockerfile.test .

# Generate a baseline (single fixture):
bun run --cwd packages/producer docker:test:update <fixture-name>

# Generate all baselines (rarely needed):
bun run --cwd packages/producer docker:test:update

The --update flag writes output/compiled.html and output/output.mp4 from the current render. Without --update, the harness compares against those baselines.

Running the harness locally

# Run every fixture (parallel, in-process mode — the default).
bun run --cwd packages/producer docker:test

# Run a single fixture:
bun run --cwd packages/producer docker:test font-variant-numeric

# Run sequentially (lower memory):
bun run --cwd packages/producer docker:test -- --sequential

Harness modes

--mode=<value> chooses which render path the harness exercises:

Mode What it calls Use for
in-process (default) executeRenderJob Day-to-day baselines. This is the same path the hyperframes render CLI takes, and it is what produced every existing output/output.mp4.
distributed-simulated plan()renderChunk() × N → assemble() from @hyperframes/producer/distributed Validates the distributed pipeline against the in-process baseline. No Temporal or Lambda involvement — the controller and chunk worker are both this process.

--mode=distributed-simulated

bun run --cwd packages/producer docker:test -- --mode=distributed-simulated
bun run --cwd packages/producer docker:test font-variant-numeric -- --mode=distributed-simulated

The distributed pipeline cannot run every fixture. Fixtures that fail any of these gates are skipped with a clear log line (and counted as passing in the summary):

  • fps.den !== 1 — distributed mode is integer-fps only (no NTSC).
  • fps.num ∉ {24, 30, 60} — closed set per DistributedRenderConfig.
  • format === "webm"plan() refuses webm.
  • hdr === true — distributed mode is SDR-only at v1.

Both modes use the fixture's authored minPsnr as the per-test threshold — distributed must clear the same quality bar in-process clears against the same frozen baseline. (Internal contract: distributed vs in-process renders of the same fixture should clear 50 dB PSNR against each other within the same Docker image. Against the frozen committed baseline, neither mode reaches that consistently due to shared encoder/JPEG-capture jitter — that's why the fixture's authored threshold gates here, not the 50 dB contract value.) An absolute 10 dB pathology floor catches fully-black-output regressions when a fixture authors a permissive threshold. A distributed failure at the fixture's own threshold means the distributed pipeline has drifted — file an issue rather than relaxing the fixture.

--update is incompatible with --mode=distributed-simulated: the in-process renderer is the source of truth for baselines, and the distributed mode's job is to verify the contract against the same baseline.

Validating PR 4.1 (the harness mode itself)

The smallest fixtures (font-variant-numeric, many-cuts) are sufficient to verify the mode plumbing end to end:

docker build -t hyperframes-producer:test -f Dockerfile.test .

# In-process: existing behavior, unchanged.
bun run --cwd packages/producer docker:test font-variant-numeric
bun run --cwd packages/producer docker:test many-cuts

# Distributed-simulated: same baselines, distributed pipeline.
bun run --cwd packages/producer docker:test font-variant-numeric -- --mode=distributed-simulated
bun run --cwd packages/producer docker:test many-cuts -- --mode=distributed-simulated

Both modes must pass at each fixture's authored minPsnr against the existing baseline. If --mode=distributed-simulated fails where --mode=in-process passes, the distributed primitive has a regression — file an issue rather than relaxing the fixture's threshold.

Distributed-only fixtures

Fixtures under tests/distributed/<name>/ are authored specifically for the distributed pipeline. They follow the same meta.json schema as the top-level fixtures, but they always set chunkSize / maxParallelChunks so a plan() over the fixture produces N>1 chunks. Each fixture exercises one of:

  • per-format chunk-boundary correctness (mp4 H.264, mp4 H.265, ProRes, png-sequence)
  • per-adapter chunk-seam state preservation (GSAP, Anime.js, Three.js, Lottie, CSS, WAAPI)

Each distributed fixture covers one or more equivalence axes — see the meta.json description field for what a given fixture is locking in.

Fixture pattern (4.2 onward)

Each tests/distributed/<name>/ fixture has the same structure as a top-level fixture (meta.json + src/index.html + output/output.mp4). Differences worth knowing:

  • renderConfig.chunkSize is required — pick a value that yields N≥2 chunks for your fixture's frame count (e.g. 60 frames at chunkSize: 15 produces N=4). Without this the fixture renders in a single chunk and never exercises the seam.
  • The fixture's ID on the CLI is just <name> (no distributed/ prefix). bun run --cwd packages/producer docker:test mp4-h264-sdr works the same as for a top-level fixture.
  • The distributed tag is informational — it doesn't gate any tag-based filter today. Add it so the fixture is easy to find by tag.
  • The composition should stress state continuity across the chunk seams: an animation crossing a seam, a counter, a rotation. A fully-static composition would pass even if chunk-boundary state was broken.
  • Baselines must be generated inside Docker — see the section above. The baseline is rendered by the in-process renderer (the source of truth for golden output); --mode=distributed-simulated is validated against the same baseline.

Tags

Common tags values control which fixtures the default bun test invocation runs. --exclude-tags transparency (the default for bun test) skips webm/png-sequence alpha fixtures that need a working chrome-headless-shell alpha pipeline.