* fix(producer): localize remote media sources + strip audio crossorigin
Two bugs affecting compositions that use remote S3 URLs for video/audio.
Bug 1 — Remote <video>/<audio> sources cause blank frames
The renderer (Puppeteer) must buffer all video elements to readyState >= 2
before frame capture begins. With 10+ large S3 clips, Chrome exhausts
pageReadyTimeout and every clip renders as a blank black frame. Fix:
localizeRemoteMediaSources() downloads all remote <video>/<audio> src
URLs in parallel during compilation and rewrites the src attributes to
local paths served by the file server, eliminating the buffering race.
Bug 2 — crossorigin on <audio> elements not stripped
htmlCompiler.ts already stripped crossorigin from <video> and <img>
(hf#1140) but missed <audio>. Compositions with crossorigin="anonymous"
on audio elements caused CORS-mode failures against the localhost file
server. Extended the strip to cover <audio>.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(producer): basename portability + localizeRemoteMediaSources tests
Addresses Rames' review on hf#1146:
- Replace `absPath.split('/').at(-1)` with `path.basename(absPath)`. On
Windows, path.join emits backslash-separated paths; split('/') returns
the whole path as a single element, producing a garbage relPath.
path.basename delegates to the OS separator on the current platform.
- Export `localizeRemoteMediaSources` for unit testing. Tests verify:
- Successful download rewrites src to _remote_media/ path
- Download failure preserves original URL without throwing
- Duplicate src URL across two tags → single fetch call (dedup)
- Local (non-HTTP) src paths are not rewritten
- Both double-quoted and single-quoted src attributes are rewritten
- basename extraction is correct on POSIX paths
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(engine): remove amix normalize=0 to fix audio on FFmpeg 4.x/6.x
amix's normalize=0 option is absent from many FFmpeg builds (e.g.
FFmpeg 4.2 on Ubuntu 20.04). When the option is not recognized, FFmpeg
fails the entire filter graph initialization, processCompositionAudio
returns success:false, and the assembled video has no audio stream.
Replace normalize=0 + weights='1...' with the amix default behavior
(normalize=true, divides by track count) and multiply the master output
gain by the track count to restore the original per-track volumes.
The net volume is identical across all FFmpeg versions.
Fixes #1136-adjacent: reported as 'audio doesn't play' in rendered MP4.
* fix(producer): strip img crossorigin + fix audioExtractor normalize=0
Two follow-up fixes:
1. htmlCompiler: strip crossorigin attribute from <img> elements during
compilation. External images (e.g. S3) with crossorigin='anonymous'
force CORS-mode requests against the renderer's localhost file server,
which S3 rejects → images render blank. Matches the existing video
strip at line 261.
2. audioExtractor: same amix normalize=0 bug as audioMixer.ts. The
audioExtractor path is used for <video data-has-audio='true'> mixing
in the CLI's local render pipeline; on FFmpeg 4.x it would also drop
audio silently. Fix: remove normalize=0, compensate with volume=N.
* test(engine,producer): pin amix normalize contract + img crossorigin strip
- audioMixer.test.ts: assert filter has no normalize=/weights=; add
3-track test confirming compensatedGain = masterGain × N = 3
- htmlCompiler.test.ts: parallel tests for img and video crossorigin
strip (covers both elements, not just video)
* test(producer): regenerate gsap-letters-render-compat baseline
b2828e48 deferred __renderReady until the root timeline is bound (May 24).
The baseline was generated May 18 under the old premature-ready behavior,
so the renderer now captures frames at slightly different animation states
for the back.out(1.8) letter stagger. 85/100 checkpoints were below the
30 dB PSNR threshold.
Regenerated in Docker with the pinned chrome-headless-shell@148.0.7778.167.
* test(producer): regenerate 7 stale regression baselines in Docker
Runtime changes since last baseline generation caused visual drift in 7 suites.
All regenerated with chrome-headless-shell@148.0.7778.167 inside Dockerfile.test.
Failures before regen:
- many-cuts: 1 frame
- overlay-montage-prod: 1 frame
- pip-video-late-host: 29 frames
- spanish-empire-cdn-inline: 74 frames
- style-18-prod: 24 frames
- style-7-prod: 50 frames
- typegpu-adapter: 75 frames
All 51 suites pass locally after this commit.
resolveChunkPlan caps chunkCount at maxParallelChunks from the naive
count, then rounds effectiveChunkSize up to ceil(totalFrames /
chunkCount). When that ceil rounds up, the first (chunkCount - 1) chunks
can already cover every frame, so buildChunkSlices emits a final slice
with startFrame >= totalFrames — an empty [n, n) or inverted range.
renderChunk rejects it (framesInChunk <= 0) and, under Step Functions
retries, fails the whole distributed render even though [0, totalFrames)
is fully covered.
This is reachable from the user-facing CLI: `hyperframes lambda render
--chunk-size 10 --max-parallel-chunks 12` on a ~4s/30fps (121-frame)
composition yields chunkCount=12, effectiveChunkSize=11, and a 12th slice
of [121, 121).
Tighten chunkCount to ceil(totalFrames / effectiveChunkSize) after the
size is finalized, so the union stays exactly [0, totalFrames) with no
empty tail. This only lowers chunkCount in the explicit-small-chunkSize
case; the auto-sized and large-chunkSize paths already satisfy
ceil(totalFrames / effectiveChunkSize) >= chunkCount, so it's a no-op
there (existing tests' chunkCount values are unchanged).
Adds a regression test for the 121/10/12 case plus a grid property test
asserting contiguous, non-empty, exact coverage across explicit sizes.
Co-authored-by: Carlos Alcaraz <193642530+calcarazgre646@users.noreply.github.com>
* fix(producer): recover from worker crashes instead of hanging the render
Both the shader-transition and png-decode-blit worker pools freed a
crashed worker's slot (busy=false, current=null) but left it in the slot
list and never marked it dead. A later run() then selected the dead slot
via slots.find(s => !s.busy) and dispatched to its terminated worker,
where postMessage is a silent no-op (no throw, no reply) — so the task
promise never settled. In the HDR hybrid capture loop, which pipelines
blends across N DOM workers and awaits every dispatch, that wedges the
whole render with no fail-fast.
The crash handlers also never drained the queue, so a queued task could
wait forever for a slot that had died.
Mark a slot dead on error/exit, exclude dead slots from dispatch and from
run()'s slot selection, and fail fast: when no live workers remain, reject
queued tasks and reject new run() calls rather than hanging. This keeps
the pools' existing no-respawn, fail-fast intent; it just actually fails
fast instead of wedging.
Adds crash-recovery tests to both pools via a fixture worker that throws
on its first message, asserting the in-flight task, queued tasks, and
subsequent run() calls all settle rather than hang.
* fix(producer): address review nits on worker-pool crash recovery
- Reword the dead-marking comments in both onWorkerError handlers: the
flag is set before rejecting and before draining the queue, not
"before anything else" (current/busy are cleared first).
- Rename the shader pool's all-slots-die test to match the png pool's
equivalent; the size-2 fixture crashes every worker, so there are no
surviving workers serving.
---------
Co-authored-by: Carlos Alcaraz <193642530+calcarazgre646@users.noreply.github.com>
* fix(studio): gsap panel bug bash — clamping, overlay, click cycling, visibility toggle
- opacity/autoAlpha clamped to [0,1] (display 0–100%) — eliminates -30%/190% edits
- `visibility` renders as a boolean toggle; only available to add in `set` tweens
- ease curve section: use aspect-ratio container so control circles are not oval
- MetricField scroll only fires when the input is focused (was triggering on scroll-over)
- preview overlay clipped to its container (overflow-hidden) — no bleed into panels
- `fromTo` method label updated to "From → To" (was "Animate", same as `to`)
- repeated click at same position cycles through stacked/overlapping elements (#1124, #1125)
resolveAllVisualDomEditTargets returns the full z-stack; subsequent same-spot clicks
advance through all selectable layers at that coordinate
- fallow-ignore-next-line complexity on pre-existing complex functions surfaced by
branching from fix/gsap-fromto-panel rather than main
Closes#1124, #1125
* fix(studio): address Vai+Rames follow-up notes on hf#1122
- extract buildTweenSummary to gsapAnimationHelpers.ts (now testable)
- add tests for all buildTweenSummary branches including fromTo
- extract requireAnimation/requireFromToAnimation helpers in files.ts,
eliminating the parse→find→guard pattern repeated across three switch
cases and removing the fallow-ignore-next-line complexity bypass
- add 400 guard: add mutation with fromProperties on non-fromTo method
now returns 400 instead of silently dropping fromProperties
- add test for the 400 guard
* fix(studio): buildTweenSummary formats percent props as 0-100% not 0-1
* fix(studio): show all .html files as compositions in sidebar
The Comps sidebar only listed index.html and files under a compositions/
subdirectory. Any other .html file in the project root was invisible and
could not be loaded as a composition preview.
Broadened the filter in useFileManager and the activeCompPath guard in
App.tsx to treat every .html file as a selectable composition.
Also excluded App.tsx from the filesize pre-commit check — the file is
already 652 lines (decomposition tracked in PR #724).
* fix(studio): detect compositions by data-composition-id, not path convention
The previous approach filtered compositions by path convention (index.html
or compositions/ subdirectory). Any .html file outside that convention was
invisible in the Comps sidebar.
The server now scans each .html file for data-composition-id and returns
a compositions[] field in the project API response. The client uses this
server-provided list instead of filtering locally. This means any .html
file that is a real HyperFrames composition shows up regardless of where
it lives in the project tree.
* fix(studio): rename Ask agent to Copy prompt to AI agent, show context preview
Updated the property panel button label from "Ask agent" to "Copy prompt
to AI agent". Updated the modal title to match. Added a collapsible
"Context included in prompt" details section to the modal that shows the
element metadata that will be included when copying.
* fix(studio): wire contextPreview to agent modal
Passes composition path, source file, selector, tag, and text content
to the AskAgentModal so the context preview section is visible.
* fix(core): seek timeline to current time after initial bind
When bindRootTimelineIfAvailable captured a GSAP timeline for the first
time, it paused it but never seeked to state.currentTime. This left
fromTo tweens stuck at their immediateRender "from" state (e.g. opacity 0)
even after the user scrubbed past the tween's end. The polling rebind
path already seeked to previousTime — the initial bind was the only path
that skipped it.
* feat(core): add gsap_timeline_not_registered lint rule
Warns when a composition creates gsap.timeline() but never registers it
in window.__timelines. Without registration, the runtime cannot discover
the timeline, and animations will not play during preview or render.
Skips the warning for sub-compositions (template-based) which inherit
the parent's timeline context.
* fix(studio): address hf#1126 review feedback
- Extract buildAgentContextPreview into domEditingAgentPrompt.ts and
import it in App.tsx, removing the inline computation that pushed
App.tsx past the 600-line CI gate
- Switch isCompositionFile from sync readFileSync to async readFile with
Promise.all, and use a regex test instead of string includes
- Move PERCENT_PROPS from AnimationCard.tsx and gsapAnimationHelpers.ts
into gsapAnimationConstants.ts (single source of truth)
- Add regression test for the totalTime initial-bind seek fix in
init.test.ts — verifies the captured timeline receives a totalTime
call on initial bind
* refactor(studio): extract App.tsx below 600 LOC, remove lefthook exemption
Extracted inspector state, studio context construction, and drag overlay
into useStudioContextValue.ts. Deduplicated block handler args via a
shared blockCtx memo. App.tsx drops from 657 to 588 lines.
Removed the App.tsx exemption from lefthook.yml — the file now passes
the 600-line gate without special-casing. Added domEditing.ts barrel to
fallowrc ignoreExports (re-exports not traceable by static analysis).
* fix(producer): cache Google Fonts woff2 per subset, preserve unicode-range
Google Fonts' css2 API returns one @font-face per (weight × unicode-range
subset) — e.g. vietnamese, latin-ext, and latin faces for the same weight,
each pointing at a distinct woff2 whose glyphs match its unicode-range.
The on-disk cache keyed woff2 files by `${weight}-${style}` only, ignoring
the subset, so every subset of a weight collided on one filename: only the
first subset in the CSS was downloaded and every later subset read it back.
For families whose CSS lists `vietnamese` first (e.g. Big Shoulders Display)
the `latin` A–Z subset was silently dropped, leaving the embedded font with
almost no Latin glyphs. The injected @font-face also omitted `unicode-range`,
so it advertised coverage it lacked and mismatched glyphs fell back to a
different font — the visible "wrong A" glitch in rendered headlines.
- Key the woff2 cache by a hash of the subset-unique woff2 URL, so each
subset is cached on its own.
- Carry each face's `unicode-range` through to the injected @font-face so
the browser selects the correct subset per codepoint (matching Google's
own CSS semantics).
- In the bundled-font Google supplement, add every subset of an uncovered
weight instead of deduping by weight (which dropped extra subsets).
- Extract per-subset download/cache into a helper to keep fetchGoogleFont
within complexity limits.
Adds a hermetic regression test (injected fetch + temp cache dir) that fails
on the old cache-by-weight behavior and passes with the per-subset cache.
* fix(producer): use atomic write for woff2 font cache (CodeQL)
Replace existsSync+writeFileSync TOCTOU pattern with try-read-first +
O_CREAT|O_EXCL (wx flag) atomic write. Eliminates the race window between
the existence check and the file creation, and prevents symlink-following
in shared temp directories (Lambda /tmp). Concurrent render processes that
race on the same cache entry now resolve gracefully via EEXIST handling.
* fix(producer): avoid os.tmpdir() taint for font cache path (CodeQL)
Replace tmpdir() call with literal "/tmp/hyperframes/fonts" for the
Lambda cache path. Lambda's /tmp is private per execution environment,
not a shared multi-user temp dir — semantically identical but breaks
CodeQL's taint tracking from os.tmpdir() to writeFileSync.
* revert: restore tmpdir() for Lambda font cache path
The hardcoded "/tmp" was a workaround for a CodeQL false positive.
Lambda's /tmp is private per execution environment; the write already
uses O_CREAT|O_EXCL + mode 0o644. Dismissed the alert as false positive
via the code-scanning API instead of warping the code.
* feat(studio): GSAP tween editing in Design panel
Add a GSAP animation editor to the studio Design panel: select an element,
view and edit its tweens (properties, easing, timing), add/delete animations,
and drag custom bezier speed curves — all persisted back to the composition
HTML. Gated behind VITE_STUDIO_ENABLE_GSAP_PANEL.
Parsing of existing GSAP source now uses a recast + Babel AST parser instead of
regex, giving scope resolution, stable tween IDs, and round-trip preservation of
extras and unresolved raw values.
recast compiles to CommonJS that calls require("fs"), which breaks browser and
Vite SSR bundles. To contain it, @hyperframes/core is split into an isomorphic
layer and a Node-only AST layer:
- gsapSerialize.ts holds the recast-free helpers (serialization, keyframe
conversion, validation, shared types). htmlParser.ts is now fully isomorphic.
- parseGsapScript and the script-mutation helpers live in gsapParser.ts,
reachable only via the @hyperframes/core/gsap-parser subpath, loaded
server-side by the studio-api mutation routes and the linter via dynamic
import (recast stays external under SSR).
- The barrel and the gsap-constants subpath are recast-free, so studio browser
bundles never trace recast.
Adds AST parser unit + stress coverage and e2e helpers for the panel.
* fix(lint): await async lintHyperframeHtml in all callers
lintHyperframeHtml became async (gsap rules use dynamic import)
but lintProject and check-hyperframe-static weren't awaiting it,
causing typecheck failures and runtime crashes in CI.
Also wire LintRule type in gsap rules to fix fallow unused-type
finding, and suppress render.ts exported-for-tests symbols.
* fix(producer): normalize error messages to prevent [object Object] in telemetry
When a render fails and the caught value is a plain object (not an Error
instance), String(error) produces [object Object], masking the real error
in PostHog telemetry (~24 errors/day).
Add normalizeErrorMessage() that tries Error.message, string passthrough,
.message on plain objects, JSON.stringify, and String() as a last resort.
Apply it on the two telemetry-feeding paths: the main render failure
handler (renderOrchestrator.ts:2099) and buildRenderErrorDetails
(cleanup.ts), plus the error classifier isRecoverableParallelCaptureError
so timeout detection works even when the thrown value is a plain object.
* fix: address review — normalize CLI telemetry path, captureCost fallback
* fix: use local normalizeErrorMessage in CLI to avoid cross-package resolution
The Vite test runner can't resolve runtime imports from @hyperframes/producer
since its exports point to dist/. Copy the utility into the CLI package and
import locally instead.
* fix(engine): use captureBeyondViewport on all CDP screenshot paths
Chrome's compositor rounds the viewport boundary inward under multi-tab
load, clipping the bottom/right edge of tall portrait compositions
(1080x1920). The explicit clip rect already constrains output to exact
composition dimensions, making the viewport-boundary pre-clip from
captureBeyondViewport:false both redundant and unreliable.
Set captureBeyondViewport:true on all three CDP screenshot call sites:
pageScreenshotCapture, captureScreenshotWithAlpha, and captureAlphaPng.
Add portrait-edge-bleed regression test: 1080x1920 grid with bright
magenta bottom rows, rendered with 4 workers. Any compositor clipping
at the bottom edge drops PSNR sharply against the golden baseline.
Closes#1009
* fix(engine): address review feedback on captureBeyondViewport
- Add backref comments on captureScreenshotWithAlpha and captureAlphaPng
pointing to pageScreenshotCapture for the rationale, so the next reader
doesn't treat the flag as unintentional copy-paste
- Note in test meta.json that the static grid fixture covers the
capture-side clipping path but not the video-element compositor surface
timing that produces the t≈37s self-healing in #1009
* test(producer): use video element in portrait-edge-bleed regression test
Replace the static CSS grid with a 1080x1920 portrait video element —
matches the original bug report shape where the compositor surface
allocation timing causes the bottom-edge clipping. The video has a dark
top region and bright magenta bottom 480px, so any viewport clipping at
the bottom edge drops PSNR sharply. Baseline regenerated in Docker with
4 workers.
* test(producer): add parallel capture regression test
Add a regression fixture that forces workers: 2, ensuring the parallel
capture code path (browser-per-worker in BeginFrame mode) is exercised
in CI. All existing fixtures pin workers: 1, so this is the first test
that would catch a regression in the multi-worker pool isolation fix
from PR #1087.
The composition is 5s @ 30fps (150 frames), which exceeds both
MIN_FRAMES_PER_WORKER * 2 (60) and minParallelFrames (120), so the
parallel coordinator will always split work across workers.
Baseline output/output.mp4 must be generated inside Dockerfile.test
before the fixture can run in CI.
* test(producer): bump parallel capture test to 4 workers
Matches realistic auto-mode worker counts (4-6 on typical machines),
not just the minimum (2) that triggers the bug.
* test(producer): add golden baseline for parallel capture regression
Generated inside Dockerfile.test on amd64 Linux (Docker image
hyperframes-producer:test) to match the CI rendering environment.
* test(producer): address review feedback on parallel capture test
- Add fixture to shard-5 in regression.yml so CI actually runs it
- Reframe description: multi-worker path coverage (frame distribution,
reorder buffer, per-worker browser lifecycle), not GPU-specific crash
guard — SwiftShader CI can't reproduce the hardware compositor race
- Remove dead @keyframes count-up (content doesn't apply to div)
- Remove unused CSS animation reference on .counter
- Regenerate golden baseline with cleaned-up HTML
* fix(producer): replace rAF + CSS keyframes with GSAP in parallel-capture test
The composition used requestAnimationFrame for a frame counter and CSS
@keyframes for animations, which triggered screenshot capture mode
(non-deterministic across workers) and caused 29 PSNR failures in CI.
All animations now use the GSAP timeline, keeping the render in
deterministic BeginFrame mode. Baseline regenerated in Docker.
Renders a 6s composition with a pip video inside a sub-composition
host starting at t=3. Verifies the video is visible during the host
window and not double-offset to t=6. Baseline generated in Docker.
- Add comment explaining hardcoded fps=30 (runtime's canonicalFps
default, not exposed on PlayerAPI)
- Add cross-reference comments between init.ts and fileServer.ts
explaining their different __renderReady timing semantics
The cfr re-encode pass hardcodes `-c:v libx264`. Pairing it with
`codec: "h265"` would silently transcode the h265 chunks to h264.
Detect the encoder discriminant in `meta/encoder.json` and throw a
typed error parallel to the existing non-mp4 format guard, so callers
surface the conflict instead of producing a wrong-codec deliverable.
— Rames Jusso
Distributed-render output today uses -c:v copy through concat → mux →
faststart, which means PTS timestamps from each chunk pass through
unchanged. Container r_frame_rate is exact (#1040 + this PR's parent),
but stream-level avg_frame_rate stays PTS-derived and can land on
fractional rationals like 27648000/921677 over a 60s render. Same for
sub-ms duration drift.
This is the achievable bar within -c copy stream-copy concat. For most
consumers (browser playback, YouTube, etc.) the difference is invisible.
For downstream tools that strict-check avg_frame_rate or
ms-precision duration (broadcast workflows, frame-accurate compositors,
some third-party transcoders), it matters.
Adds an opt-in cfr config flag (default false). When true, the
assemble step's final pass re-encodes with -fps_mode cfr -r <fps>
instead of -c copy, producing exact CFR output. Trade-off: ~2-5x the
stitch time for a 60s 1080p clip; second-generation H.264 quality loss
is negligible at -crf 18 but is non-zero.
The v0.6.39 fix added -r <fps> to the multi-chunk concat ffmpeg
invocation but didn't reach the single-chunk pass-through path,
which is taken when totalFrames * fpsDen / fpsNum fits in one
chunk. Result: 1-chunk renders shipped with fractional
r_frame_rate (e.g. 359/12) while multi-chunk renders shipped
with exact 30/1.
Single-chunk path now goes through the same -r <fps> + -c copy
ffmpeg invocation as the concat path, ensuring uniform exact
r_frame_rate metadata across all chunk-count configurations.
Adds a regression test exercising the 1-chunk path and asserting
r_frame_rate === "<fpsNum>/<fpsDen>" exact.
- Always log when html-in-canvas pins to 1 worker, not just on
explicit --workers override (Vance, James)
- Assert recommendScreenshot is true for htmlInCanvas detection —
pins the load-bearing coupling with the screenshot paint-force
protocol (Vance, Magi)
- Skip capture calibration for htmlInCanvas auto-sized renders to
avoid wasting 600ms–3s on an estimate that gets thrown away (Vance)
- Add TODO documenting Chrome's two root-cause mechanisms: paint
cache race and SwiftShader contention (James, Magi)
Chrome's drawElementImage API does not support concurrent usage across
multiple browser instances — running >1 capture worker causes flickering
artifacts. Detect the layoutsubtree canvas attribute during compilation
and unconditionally pin workers to 1, overriding both auto-sizing and
explicit --workers flags.