* docs(hyperframes): route /slideshow as a workflow, not a capability
Move /slideshow out of the domain-skill capability map and into the
intent router as a top-level workflow. It is an intent-gated orchestration
that produces a navigable deck, not an atomic capability loaded on demand.
Also clarify that workflows need not output a video: /slideshow builds a
deck and /remotion-to-hyperframes ports a composition. Adds the cheat-sheet
row, a disambiguation bullet, and the per-workflow detail block.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(hyperframes): sharpen /slideshow disambiguation to match its intent gate
Reword the slideshow disambiguation bullet to match the skill's actual
intent-confirmation behavior: an explicit "slideshow" request proceeds
directly; an adjacent trigger ("deck / slides / presentation / convert this
page") makes /slideshow confirm before authoring and switch to the
appropriate non-slideshow workflow if not. Drops the inaccurate "may
actually want a video" narrowing.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: handle caption skin workflow
* docs(skills): simplify the finalize step across video workflows
- Drop --strict-layout; all skills use plain `hyperframes inspect`
- Add the caption text_box_overflow false-positive note to faceless-explainer
- On a failed check, the orchestrator makes the cheapest safe edit itself
(no worker re-dispatch / Step 3 backtrack language)
- Snapshot: glance at the stitched contact-sheet.jpg and move on
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(auth): onboarding-first `auth status` + shared TTS/BGM preflight
When no HeyGen credential is configured, `hyperframes auth status` now
prints registration-first guidance instead of a terse error:
- Interactive / agent-driven sessions get sign-in guidance led by
`hyperframes auth login` (the OAuth step that also creates an account
and is shared with heygen-cli), and never steer users to a per-repo
`.env`. CI / non-interactive runs get a terse note. Exit 1 is kept so
the "am I logged in?" `$?` contract still holds.
- It probes which local engine voice/music will fall back to (Kokoro /
MusicGen, mirroring the skill resolution order) and whether their
Python deps are installed, with a pip hint when missing. `--json`
exposes `recommended_action` + `offline_engines` for skills to branch.
- `doctor` gains matching "TTS (Kokoro)" / "BGM (MusicGen)" checks via
the same shared probe (findPython/hasPythonModules extracted to
tts/python.ts; provider resolution in audio/providers.ts).
Every TTS/BGM workflow now relays this at Step 0 (setup) instead of
improvising its own "missing key" prompt: pr-to-video, product-launch-
video, faceless-explainer, website-to-video, music-to-video. The
canonical behavior + key-priority table live once in hyperframes-media.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(pr-to-video): scale recommended video length to PR change size
Step 0 led with a fixed ~60-90s length default. Now the recommended
length is derived from the PR's diff stat (lines added+deleted, nudged
by file count) on a tier scale (trivial ~20-40s → large ~110-180s, hard
cap ~3 min), reusing the same PR peek already done to infer the angle.
The agent states the basis when proposing it, and a huge PR with one
headline change still stays tight. User can always override.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(captions): embed brand fonts whose files use separators
brandFontFaces() matched font files by stripping only whitespace, so an
underscore/hyphen-named file (TT_Norms_Pro_Bold.woff2) never matched the
family key "ttnormspro" — captions shipped with no @font-face, the
font_family_without_font_face bug. Now both family and filename normalize
away all non-alphanumerics; families match longest-key-first so a parent
family can't swallow a more specific one's files (TT Norms Pro vs Mono);
each file is claimed once; "demibold" ranks before "bold"; and when
nothing matches it warns loudly at build time instead of returning "".
Also: parseFonts() falls back to h1/h2/title/hero display roles, and the
frame-worker + caption authoring docs spell out that only shipped font
files render — no system CJK/Devanagari families on the headless renderer.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(hyperframes-media): enforce sign-in preflight on standalone BGM/TTS
A one-off "generate me a BGM" request went straight to local MusicGen
without recommending sign-in: bgm.md/tts.md framed the no-credential path
as an automatic fallback, so the generation path bypassed the Preflight
stop, and the preflight used a bare `hyperframes auth status` that isn't
on PATH in a fresh `npx skills` project.
- Preflight now applies to one-off generation as well as workflows, uses
`npx hyperframes auth status`, and says: if the CLI can't run, still
recommend signing in and STOP — never treat "no credential" as a silent
green light for local generation.
- bgm.md and tts.md point at the Preflight before generating, reframing
local generation as the fallback the user opts into, not a default.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(auth): add Authentication & API keys guide
Document signing in, the keys each capability (voice, music, capture)
uses, their resolution priority, and the fully local fallback. Add the
guide to the nav and cross-link it from the cloud deploy note and the
CLI env-var reference.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(lint): strip HTML comments in a fixpoint loop (CodeQL)
Single-pass <!-- --> removal can re-form a complete comment from
adjacent markers (e.g. `<<!-- -->!-- ... -->`), letting a decoy
<template> survive and hijack the template-boundary match. Loop to a
fixpoint, mirroring the captions.mjs precedent; add a regression test
that fails on single-pass (2 root findings) and passes on the loop.
Also wrap the build-frame.mjs node:fs imports to satisfy oxfmt — the
new copyFileSync import pushed the line past the width limit, which
was the sole cause of the Format / Preflight CI failures.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(lint): strip HTML comments with a linear scan (CodeQL ReDoS)
The fixpoint loop still ran a /<!--[\s\S]*?-->/ regex per pass, which
backtracks O(n^2) on inputs with many unterminated "<!--" — CodeQL
js/polynomial-redos (high). Looping the same regex (the prescribed
fix) never addressed this; only the regex itself does.
Replace it with an indexOf-based linear strip in utils.ts
(stripHtmlComments), kept in a fixpoint loop so markers that re-form
when a comment is removed are still stripped. 200k unterminated
"<!--" now strips in ~3ms instead of quadratic time; behavior is
otherwise unchanged — unterminated comments are kept verbatim, as the
old regex left them. The re-forming regression test still guards it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(auth): make TTS/BGM sign-in guidance accurate and runnable
From team review of the not-signed-in onboarding:
- OAuth is a `hyperframes auth login` feature only. The separate `heygen`
CLI is API-key-only — `heygen auth login` stores a pasted key, it is not
OAuth and does not create an account. Stop presenting the two CLIs as the
same OAuth/sign-up step.
- Use `npx hyperframes` in every imperative and runtime hint. Bare
`hyperframes` is not on PATH on a fresh machine (command not found); only
`npx hyperframes` is guaranteed. Also updates the JSON recommended_action.
- Drop `heygen auth login` from the terminal/skill onboarding: it needs its
own install and there is no `npx heygen`, so it was a command-not-found
trap. The shared-credential fact stays in the reference docs.
Covers the `auth status` guidance + tests, the Authentication docs, the
shared hyperframes-media preflight (SKILL, requirements, tts, error hints),
and the `npx hyperframes auth status` preflight in every TTS/BGM workflow
(pr-to-video, product-launch-video, faceless-explainer, website-to-video,
music-to-video).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
## What
Studio preview now resolves `<video src="../../assets/x.mp4">` (and the
same shape for `<img>`, `<audio>`, inline `style` `url()`, and `<style>`
CSS `url()`) against the sub-composition's URL — matching what the
server-side bundler already does for the render path.
## Why
Authored compositions live at `compositions/frames/*.html` and reference
project-root assets either as plain `assets/x.mp4` (already correct
because the main document's `<base href>` points at the project preview
root) or as `../../assets/x.mp4` (the explicit project-root-relative
form). The server-side `inlineSubCompositions` flattens sub-comps into
`index.html` and rewrites the `../`-form against the sub-comp's source
path so it resolves against the project root in the baked render.
The browser-side runtime that mounts external sub-compositions via
`fetch` did no such rewriting. So `<video src="../../assets/x.mp4">`
authored inside a `compositions/frames/scene.html` resolved against the
main document's base href, climbed above the project root, and 404'd in
Studio preview — even though the same path rendered correctly in the
final video. An OSS user (Miao Yang) hit this in a real project.
## How
Added `rewriteSubCompositionAssetPaths` to the runtime
`compositionLoader`. After parsing the fetched sub-composition HTML and
before extracting any nodes, walk the parsed document and rewrite the
same surface the server-side path touches:
- `[src]` and `[href]` attributes on every element
- `[style]` attribute `url(...)` references
- `<style>` element CSS `url(...)` references
The rewrite mirrors the producer's semantics exactly: only values that
start with `../` (or are literal `..`) are rewritten — against the
sub-composition's URL via `new URL(value, compositionUrl)`. Absolute
URLs, root-relative paths, `data:`, hash refs, and plain
`assets/x.mp4` are left untouched. **Plain relative paths must not be
rewritten** because the main document's `<base href>` already covers
them; rewriting would double-prefix the URL.
The walk recurses into `<template>` content because authored
compositions typically wrap their rendered body in a `<template>` and
`querySelectorAll` does not enter template content (it lives in a
detached `DocumentFragment`).
## Test plan
- [x] Unit tests added (6 new tests in
`compositionLoader.test.ts`): rewrites `../`-traversing src on
template-wrapped sub-comps; leaves plain relative paths untouched (no
double-prefix); leaves absolute / data / hash / root-relative URLs
untouched; rewrites CSS `url()` in `<style>` blocks and inline
`style` attributes; rewrites for non-template (full-HTML-doc)
sub-comps.
- [x] Full core test suite green (2065 tests).
- [x] Full studio test suite green (1148 tests).
- [x] Manual verification with the reporter's actual project:
before the fix one `<video>` with a `../../assets/...` src returned
`MEDIA_ELEMENT_ERROR: Format error`; after the fix all 7 `<video>`
elements load (`readyState=4`, correct `currentSrc`). The 6 plain
`assets/...` paths are *unchanged* (no double-prefix) and continue
to resolve via `<base href>` as before.
- [x] `bun run lint`, `bun run format:check`, `bun run typecheck`,
`fallow audit` all green.
Reported by Miao Yang.
— Jerrai (https://claude.com/claude-code)
* feat(telemetry): attribute renders to the authoring workflow skill
Add an optional `--skill` flag to `hyperframes render` and tag the
`render_complete` / `render_error` events with `authoring_skill`, so render
usage can be broken down per authoring workflow. The value is slug-gated (a
malformed value is ignored) and the existing anonymous / opt-out telemetry
pipeline is otherwise unchanged.
Each end-user workflow that renders now passes `--skill=<name>` on its render
command: embedded-captions, faceless-explainer, graphic-overlays,
motion-graphics, music-to-video, pr-to-video, product-launch-video,
remotion-to-hyperframes, website-to-video.
Not instrumented, by design: general-video renders freeform with no canonical
render command to attach to, and slideshow produces an interactive deck rather
than a rendered video. Both can follow up if per-skill numbers are wanted.
* fix(telemetry): address review — shared slug util, equals-form flag, invalid-value warning
- Extract the SKILL_SLUG regex + a normalizeSkillSlug() helper into
telemetry/skill.ts, shared by the `events` and `render` commands (the regex
was duplicated). `render` adopts normalizeSkillSlug (so it now trims the value,
matching `events`); `events` references the shared SKILL_SLUG. + unit test.
- `render` warns on a non-empty but invalid --skill value (e.g. a camelCase
typo) so attribution isn't silently lost — stderr only, never fails the render.
- embedded-captions render script: `--skill embedded-captions` -> `--skill=embedded-captions`.
On an older CLI that does not declare --skill, the space form leaks the value
as a positional and clobbers the project dir (resolveProject fails); the equals
form is parsed as a self-delimiting flag and safely ignored. Verified via Node
parseArgs(strict:false).
Addresses review feedback on the PR (shared util + .trim drift, version-skew
safety, invalid-value visibility).
---------
Co-authored-by: kiritowoo <295860553+kiritowoo@users.noreply.github.com>
* fix(producer): retry probe stage on transient browser errors (#1687)
The distributed render plan stage crashes when headless Chrome encounters
a transient frame detachment ("Navigating frame was detached") during
browser probe, with no retry logic. The plan tarball is never uploaded,
and all downstream chunk workers fail with S3 404.
Add a retry-with-fresh-session mechanism to the probe stage:
- `isTransientBrowserError()` classifier in the engine identifies 9
known transient Puppeteer/Chrome errors (frame detached, target closed,
session closed, protocol error, page crashed, execution context
destroyed, etc.).
- `runProbeStage()` wraps browser session creation + initialization in a
retry loop (max 2 attempts). On transient error: logs structured
diagnostics (attempt, isTransient, error message, elapsed time), closes
the crashed session cleanly, creates a fresh browser, and retries. Non-
transient errors throw immediately without consuming retry budget.
- 17 unit tests for the error classifier, 3 integration tests for retry
behavior (successful retry, immediate throw on non-transient, exhaust
retry budget on persistent transient).
Closes#1687
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: address review feedback — widen retry scope, deduplicate patterns
- Move createCaptureSession inside the retry try/catch so browser launch
failures (Failed to launch the browser process, ECONNREFUSED) are also
retried — not just initializeSession errors.
- Deduplicate transient error patterns: remove "Protocol error.*Target
closed" (subsumed by "Target closed") and "Navigation failed because
browser has disconnected" (subsumed by "browser has disconnected").
- Add browser launch failure patterns: "Failed to launch the browser
process" and "ECONNREFUSED".
- Add test for createCaptureSession transient throw (browser launch retry).
- Update test mock comment to document sync requirement with engine
pattern list.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
When the root element lacks an explicit data-duration attribute and
there is no GSAP timeline, getDeclaredDuration now computes
max(data-start + data-duration) across all sub-compositions instead
of returning zero.
getTrackStyle() can receive a falsy tag at runtime (e.g. empty string
from timeline element defaults), causing toLowerCase() and startsWith()
to throw. Default to "div" when tag is falsy.
- Platform-specific install instructions for Linux (apt) and Windows
(releases page / cmake) instead of a generic "see GitHub" fallback
- Export getInstallInstructions so doctor can reuse it
- Add whisper-cpp check to `hyperframes doctor` after the Environment
check — reports path when found, shows install hint when missing
`hyperframes init` only installed AI coding skills on the interactive path
(behind a clack confirm). When an agent drives it non-interactively (no TTY),
it just printed `npx skills add ...` and returned — so skills were never
installed and the agent later hit `Unknown skill: <workflow>`.
- init: both interactive and non-interactive branches now run
`npx skills add` with `cwd` set to the new project dir so skills land
there, not in the caller's working directory. Non-interactive additionally
passes `--yes`; when Claude Code is driving (CLAUDECODE env var), adds
`--agent claude-code` so skills target `.claude/skills/`.
- skills: `runSkillsAdd` accepts `cwd` and `extraArgs` so callers can
control where and how skills are installed.
- templates: CLAUDE.md / AGENTS.md now tell agents to run
`npx skills add heygen-com/hyperframes` to install or update skills.
Co-authored-by: Wenbo Zhu <295860553+kiritowoo@users.noreply.github.com>
* feat(skills): add bgm-to-video skill
Add the music-to-video skill: turns a music/BGM track into a kinetic
typography video. Includes the director/builder/music-reader/finalize
agents, reference contracts, beatgrid analysis script, motion-primitive
library, and starter templates.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(lint): catch CSS↔GSAP transform conflicts in scoped selectors and frame sub-compositions
gsap_css_transform_conflict existed but missed the most common real-world
shape (a label centered with CSS translateX(-50%) plus a GSAP xPercent that
stacks to -100% in the capture path), for three independent reasons:
- selector matching was exact-string, so a scoped/grouped GSAP selector
("#root .label, #root .sub") never matched a CSS class rule (.label)
- the acorn parser only captures timeline-rooted calls (tl.to/tl.set), so a
standalone gsap.set("#root .label", { xPercent: -50 }) was invisible to it
- lintProject read compositions/ non-recursively, so per-frame compositions
in compositions/frames/*.html were never linted at all
Fix: token-decompose grouped/descendant/compound selectors and match by
id/class against CSS transform rules; additionally scan standalone gsap.*
transform calls; and recurse into compositions/ subdirectories so frame
sub-compositions are linted.
Adds unit tests (grouped gsap.set repro, descendant tl.to, negative case) and
an end-to-end lintProject test that writes compositions/frames/04-*.html and
asserts the conflict is reported there.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(skills): add beat-synced montage authoring recipe
* feat(skills): unify bgm-to-video flows into music-to-video
Replace bgm-to-video, bgm-to-video-new, bgm-to-video-refactor, and the
standalone beat-sync/montage skills with a single music-to-video skill.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(skills): register music-to-video in the hyperframes router
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(skills): add music-source brief to music-to-video Step 0
Check for user-supplied audio first; otherwise guide BGM generation
via /hyperframes-media. Note the skill targets fast, high-energy BGM.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(producer): restore css-var-fonts regression baseline
Accidentally deleted by a prior `git add -A`; it is the golden output.mp4
the distributed regression harness diffs against. Restored byte-identical
to main.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* style(skills): apply oxfmt to music-to-video and router docs
Fixes the Format / Preflight CI checks on the new skill files.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(producer): store css-var-fonts baseline as raw binary, not LFS pointer
The previous restore was re-filtered into a 130-byte LFS pointer by the
.gitattributes lfs rule; main stores this fixture as a raw binary blob
committed directly. Commit the exact blob so the regression harness reads
real frames and the file matches main.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(lint): keep the fallow audit gate green
Extract rootClassStyledSelectors so the subcomposition_root_styled_by_class
rule drops below the complexity threshold, and ignore the music-to-video
reference HTML (template + motion-primitive materials forked by path, not
import-graph reachable) — same treatment as motion-graphics/grounding.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: unblock music video ci checks
* docs: refine music-to-video planning catalogs
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: e-jung <8334081+e-jung@users.noreply.github.com>
Extract rootClassStyledSelectors so the subcomposition_root_styled_by_class
rule drops below the complexity threshold, and ignore the music-to-video
reference HTML (template + motion-primitive materials forked by path, not
import-graph reachable) — same treatment as motion-graphics/grounding.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The previous restore was re-filtered into a 130-byte LFS pointer by the
.gitattributes lfs rule; main stores this fixture as a raw binary blob
committed directly. Commit the exact blob so the regression harness reads
real frames and the file matches main.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Accidentally deleted by a prior `git add -A`; it is the golden output.mp4
the distributed regression harness diffs against. Restored byte-identical
to main.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Check for user-supplied audio first; otherwise guide BGM generation
via /hyperframes-media. Note the skill targets fast, high-energy BGM.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Replace bgm-to-video, bgm-to-video-new, bgm-to-video-refactor, and the
standalone beat-sync/montage skills with a single music-to-video skill.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
gsap_css_transform_conflict existed but missed the most common real-world
shape (a label centered with CSS translateX(-50%) plus a GSAP xPercent that
stacks to -100% in the capture path), for three independent reasons:
- selector matching was exact-string, so a scoped/grouped GSAP selector
("#root .label, #root .sub") never matched a CSS class rule (.label)
- the acorn parser only captures timeline-rooted calls (tl.to/tl.set), so a
standalone gsap.set("#root .label", { xPercent: -50 }) was invisible to it
- lintProject read compositions/ non-recursively, so per-frame compositions
in compositions/frames/*.html were never linted at all
Fix: token-decompose grouped/descendant/compound selectors and match by
id/class against CSS transform rules; additionally scan standalone gsap.*
transform calls; and recurse into compositions/ subdirectories so frame
sub-compositions are linted.
Adds unit tests (grouped gsap.set repro, descendant tl.to, negative case) and
an end-to-end lintProject test that writes compositions/frames/04-*.html and
asserts the conflict is reported there.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add the music-to-video skill: turns a music/BGM track into a kinetic
typography video. Includes the director/builder/music-reader/finalize
agents, reference contracts, beatgrid analysis script, motion-primitive
library, and starter templates.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add a routing rule to the /hyperframes entry skill: after picking a
workflow, if its skill isn't available to the agent, tell the user to
install it rather than silently falling back to a guess. Covers the
targeted install (`--skill <name>`) and the install-everything one-shot
(`--all`), then re-read the workflow's skill and continue.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
## What
Two small render-side improvements for video-heavy compositions:
1. **`packages/core/src/runtime/media.ts`** — gate the per-tick `el.currentTime = relTime` set + the `el.load()` drift-recovery retry on the *absence* of a `<img id="__render_frame_<id>__">` sibling (i.e., we're in render mode + this video's visual is bypassed by frame injection + its audio is mixed by ffmpeg from source files).
2. **`packages/engine/src/services/videoFrameInjector.ts`** — probe `window.__hfReseekGpu` and `window.__hf.colorGrading.redraw` once at the first injector call; cache the booleans; skip the per-frame `page.evaluate` round-trips when neither capability is registered.
## Why
#### media.ts
During render the runtime calls `el.currentTime = relTime` on every active video per sync tick. For frame-injected videos that's pure waste:
- The visual comes from the `<img id="__render_frame_<id>__">` sibling injected by the producer's `videoFrameInjector` — the `<video>` element is `visibility: hidden`.
- Audio is mixed by ffmpeg from the source files in `runAudioStage` (separate stage) — it never goes through the in-browser audio pipeline during render.
So every per-tick seek just kicks Chrome's media pipeline (buffering checks, range fetches, decoder state changes) for no visible or audible benefit. On a 30 × 32 MB synth comp, that's ~2,400 wasted seeks per render — and the cost wasn't on the JS critical path, so it didn't show up in `avgBeforeCapture` directly. It bled into the BeginFrame compositor's per-frame screenshot time.
Preview is unaffected: the injection sibling only exists during render. In preview `hasInjectionSibling` is always false → existing seek path runs unchanged.
#### videoFrameInjector.ts
The injector hook ran `__hfReseekGpu` and `redrawRuntimeColorGrading` via `page.evaluate` on every render frame. For comps that don't register either capability (the common case — anything without WebGL/WebGPU video sub-comps or a color-grading layer), each was a no-op page-side function preceded by a ~CDP-round-trip-worth of overhead. Probing once and caching `false` eliminates that for the rest of the render.
## How was this validated
Stress shape: `synth-30-heavy` — 30 × 32 MB MP4 / 3 s each, sequenced end-to-end over a 90 s timeline (`data-composition-id` root + per-video `<video id="vid-NN" data-start data-duration data-track-index>`). Host: 8-core / 30 GB Linux.
N=3 baseline against stock `origin/main` (post-#1630), N=3 with-fix on the same machine, same corpus, fresh worker pool each run. Phase timings via `[Render:trace]` JSON; per-frame sub-breakdown via a one-line `[CapturePerf]` stderr emit (kept locally, not in this PR — `dedupPerfs` already carries the data, this branch surfaces it).
| | Baseline N=3 | With-fix N=3 | Δ |
|---|---|---|---|
| wall mean | 119.5 s ± 1.4 s | **117.3 s ± 0.9 s** | **-2.2 s (-1.8%)** |
| avg screenshot / frame | 50.0 ms | **49.0 ms** | -2.0% |
| avg beforeCapture / frame | 13.0 ms | **12.1 ms** | -7.0% |
| avg total / frame | 66.0 ms | 63.9 ms | -3.2% |
| output md5 | `5a22be64...` | identical ×3 | ✓ |
The 1 ms screenshot drop is the load-bearing signal: it confirms the kicked Chrome media-pipeline work *was* bleeding into BeginFrame compositor time, even though it wasn't on the JS critical path. Per-frame budget improved 2.1 ms × 2700 / 3 workers ≈ 1.9 s of `capture_disk` savings, which matches the observed wall delta.
This stacks cleanly with #1630 (which removed the injector's fileServer contention). #1630 moved the injector's PNG fetches off the fileServer's hot path; this PR keeps Chrome's media pipeline quiet during render so the BeginFrame compositor runs unhindered.
## Test plan
- [x] Local-CLI render on `synth-30-heavy` × N=3 baseline + N=3 with-fix; wall, per-frame, md5 captured (above).
- [x] Lint / format / typecheck via lefthook pre-commit (`oxlint`, `oxfmt`, `fallow audit`, `tsc --noEmit` across `@hyperframes/core` + `@hyperframes/engine` + `@hyperframes/producer`).
- [ ] *Real-world video-heavy comp validation* — would love a Magi / Miga eye on a HF-heygen-stripe-shape or a Rahino-shape comp to confirm there's no audible artifact on unmuted videos. The change shouldn't affect them — in render mode the audio path is ffmpeg, not the in-browser pipeline — but a sanity-check render is cheap.
## Scope notes
- *Not addressed in this PR*: the user-facing request for an upfront-extract concurrency cap (`Promise.all` in `extractAllVideoFrames` is currently unbounded across all videos). Filing as a follow-up PR — different layer of the pipeline, different user surface (CLI flag), worth keeping separate for review.
- *Edge case*: in the calibration test-frame phase, the injection sibling may not yet exist when drift recovery first checks a video at the very start of its active window. The gate correctly defaults to "no sibling → run the seek" in that case, which is the existing behavior.
_Authored by Jerrai (Rames team)._
* fix(compiler): skip CSS var() in font resolver — fixes FONT_FETCH_FAILED on distributed renders
The font scanner treated `var(--ui-font)` as a literal font family name,
causing fail-closed distributed renders to throw FONT_FETCH_FAILED for
any composition using CSS custom properties in font-family declarations.
CSS var() expressions resolve at browser paint time, not at compile time.
The regex-based font scanner cannot resolve them statically — skip them
and let headless Chrome handle variable substitution during render.
Closes#1654
— Miga
* test(regression): add distributed css-var-fonts fixture
Regression test for compositions that use CSS custom properties in
font-family declarations. Exercises the var() skip guard in
extractRequestedFontFamilies() under the distributed renderer's
fail-closed font resolution path.
Baseline needs to be generated on first CI run with --update.
— Miga
* fix(compiler): address review feedback — mixed declaration test + validator TODO
Add unit test verifying concrete fonts alongside var() in mixed
declarations still get resolved (non-aggression pin).
Add TODO(#1654) in validateNoSystemFonts for the var()-as-primary gap
flagged by both reviewers.
— Miga
* fix(test): correct stale 4xx fail-closed test expectations
The 4xx tests expected no throw, but that was the contract before #1255
added the system font capture path (Path 3). Post-#1255, a font that
gets 4xx from Google Fonts AND isn't a bundled alias AND has no system
font IS genuinely unresolvable — fail-closed mode should throw.
The 4xx distinction still matters at the fetch level (no retry, treated
as deterministic "not served"), but at the final unresolved check, a
completely unresolvable font must throw regardless of the HTTP status
that caused the Google Fonts path to return empty.
Updated tests to match the actual contract: 4xx + unresolvable = throw.
Also set allowSystemFontCapture: false to match how distributed renders
(plan.ts:799) actually call the function.
— Miga
---------
Co-authored-by: Miguel Ángel <miguel.sierra@heygen.com>
Probes the rendered output for video and audio stream durations after
render and fails the test if they differ by more than 0.5s. Catches
mux-level truncation regressions like the ffmpeg -shortest bug (#1648)
where one stream gets silently cut short.
Runs on all non-png-sequence fixtures with audio — no new meta.json
field needed since this is a universal invariant, not a per-fixture
threshold.
FFmpeg 6.0 (bundled by ffmpeg-static) has a regression where -shortest
combined with -c:v copy over-truncates the video stream while leaving
audio untouched. The flag is also redundant — the audio mixer already
pads/caps all tracks to totalDuration via apad=whole_dur and -t.
Closes#1648
- include target in useGsapAnimationsForElement fetch key so a selection
change triggers a re-fetch even at the same cache version
- add gsapCacheVersion to useDomEditPreviewSync deps so the selection
re-syncs after every soft reload
- trim manualEditsDom.ts to 600 LOC (filesize compliance)
Fixes#1645
* fix(producer): inline base64 frames in injector to unblock video-heavy renders
The URL-served frame path (PR #596) hands each injected `<img>` a fileServer URL
instead of a base64 data URI, on the theory that shipping a short URL through
`page.evaluate` beats shipping a multi-MB base64 string per frame. That holds
when the fileServer is otherwise idle.
But on video-heavy compositions, the same fileServer also serves every
`<video>.src`. The runtime's drift-recovery branch (`runtime/media.ts:294-302`)
issues `el.load()` on the underlying `<video>` during seeks, kicking off
full-file downloads that occupy the fileServer's single Node event loop (it
uses `readFileSync` and offers no `Accept-Ranges`). The injector's
`<img>.decode()` then queues behind those video fetches and is never serviced
before puppeteer's protocol timeout fires, surfacing as
`Runtime.callFunctionOn timed out` in `capture_streaming`.
Reproducer (30 × 32 MB videos / 90 s comp / 8-core / 30 GB host):
baseline (broken corpus) 537 s render fails
baseline (corpus-fixed) 428 s render fails
this fix (drop frameSrcResolver) 121 s render succeeds, 69 MB MP4
Control corpus (30 × 1.6 MB / 60 s) shows no regression: 137 s with this
change vs ~135 s on \`main\`. The \`createCompiledFrameSrcResolver\` builder and
the \`frameSrcResolver\` option stay in the codebase, just unused for now —
re-enabling them behind a proper gate ("only use URL-served frames when the
page has zero fileServer-bound \`<video>.src\` traffic") is a follow-up. The
cache memory ceiling (\`frameDataUriCacheBytesLimitMb\`, default 1500 MB above
8 GB hosts) already bounds the cost of base64 inlining.
— Jerrai
* refactor(producer): drop unused frameSrcResolver builder import in render orchestrator
Followup to the previous commit. The void-call and the
`createCompiledFrameSrcResolver` import in `renderOrchestrator.ts` were left
behind as a no-op breadcrumb for the future gating PR. Code review (PR #1630)
correctly flagged this as dead code — the builder is a pure factory with no
side effects, so calling it and discarding the result is just wasted CPU.
Remove both and explain in the in-source comment where the builder still
lives, so the gating PR knows where to re-import from.
— Jerrai
* fix(core): register sub-composition timelines after async build + lint rule
When a composition builds its GSAP timeline inside document.fonts.ready (or any
async callback), registering window.__timelines[id] BEFORE the build leaves an
EMPTY timeline registered. The runtime's sub-composition readiness gate treats
"key present" as "ready" and nests the child once — an empty timeline gets
nested empty and is never re-nested, so the frame renders blank when used as a
sub-composition.
- registry/blocks/code-{diff,highlight,morph,scroll,typing}: register the
timeline AFTER the fonts.ready build completes, then call
window.__hfForceTimelineRebind() to re-nest now that it is populated.
- core lint: add rule gsap_timeline_registered_before_async_build to flag the
early-registration anti-pattern, with tests.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(studio): import commitGsapPositionFromDrag from its actual module
The function was split out into gsapDragPositionCommit.ts in #1605, but the
test kept importing it from ./gsapDragCommit, which no longer exports it —
yielding 'is not a function' at runtime. Import from the correct module.
Inherited main breakage (same fix as #1631/#1635); fixes the Test CI check on
this branch independently of merge order.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(core): escape digit-leading id selectors in standalone sub-composition preview
A CSS identifier cannot start with a digit, so an authored rule like
`#01-wall-pushes-back { ... }` is an invalid selector and the browser drops
the whole rule — taking the root's size/background with it. A full
composition masks this (the host stretches/paints the frame), but a
standalone preview has no host, so the root collapses to height:0 +
transparent and renders blank.
extractFullDocumentParts now rewrites `#<digit-leading-id>` selectors to
their escaped valid form (`#\30 1-...`, still matching the element id),
scoped to ids actually present and matched only as `#id` not followed by an
ident char so hex colors are never touched. Also harden the <template>
inner-HTML extraction to use the DOM instead of a greedy regex. Tests added.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit e43b377fc0)
* feat(cli): surface captured video clips in asset descriptions
generateAssetDescriptions now reads extracted/video-manifest.json and emits
each downloaded clip first, tagged [video], with its DOM heading/caption and
dimensions — motion clips are usually the strongest hero material and
downstream planners key off the [video] marker.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(studio): import commitGsapPositionFromDrag from its actual module
The function was split out into gsapDragPositionCommit.ts in #1605, but
the test kept importing it from ./gsapDragCommit, which no longer exports
it — yielding `is not a function` at runtime. Import from the correct
module to match the production import in gsapRuntimeBridge.ts.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(core): address review nits on standalone sub-composition preview
Review follow-ups (#1631), all non-blocking polish:
- contentExtractor: use path.basename() instead of localPath.split('/').pop()
so video filenames resolve correctly on Windows-style paths too.
- subComposition: document that only the leading digit needs CSS escaping
(CSS Syntax L3 §4.3.11) on escapeLeadingDigitIdent.
- tests: pin three previously-uncovered paths — multiple digit-leading ids in
one composition, a digit-leading id inside compound/combinator selectors, and
the promoteTemplateCompositionId no-op when the <template> has no id.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(product-launch-video): restructure onto script-driven architecture
Move product-launch-video onto the shared script-driven authoring flow:
build-frame remixes a hyperframes-creative preset onto brand tokens, audio
routes through the shared hyperframes-media engine, per-preset caption skins,
and every frame is authored as a directed shot. Removes the old bespoke
scripts (captions/validate/prep/hoist/…) in favour of the shared lib.
assemble-index.mjs keeps upstream #1629's blank/partial scene-file guard
(reject an empty or markup-less scene file at assembly, before emitting
data-composition-src, and re-dispatch) carried onto the restructured reader.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(pr-to-video): restructure onto script-driven architecture
Move pr-to-video onto the shared script-driven authoring flow: ingest.mjs
folds the gh PR artifacts into the synthetic capture package the shared
backend (build-frame / captions / assemble-index) reads, add the mechanism
beat, route audio through hyperframes-media, and remix a hyperframes-creative
preset onto brand tokens via the shared lib.
- Fix skill name: pr-to-video-refactor -> pr-to-video (match directory).
- Drop a stale faceless-explainer-refactor reference in an ingest.mjs comment.
- assemble-index.mjs keeps upstream #1629's blank/partial scene-file guard.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(faceless-explainer): restructure onto script-driven architecture
Move faceless-explainer onto the shared script-driven authoring flow:
every visual is invented (typography / abstract graphics / diagram / data-viz)
and authored through the shared backend (build-frame remixes a
hyperframes-creative preset onto tokens, audio via hyperframes-media,
assemble-index builds the standalone index.html) using the shared lib.
- Fix skill name: faceless-explainer-refactor -> faceless-explainer (match directory).
- assemble-index.mjs keeps upstream #1629's blank/partial scene-file guard.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(skills): refresh test-skills-fresh.sh workflow roster
Update the install-and-verify harness to the current surface: 10 workflows
(adds website-to-video, embedded-captions, graphic-overlays, slideshow;
drops the removed footage-recut) and refreshed example prompts.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* style(product-launch-video): oxfmt storyboard.mjs
Run oxfmt over lib/storyboard.mjs — formatting only, no logic change.
Fixes the Format / Preflight CI check.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(studio): import commitGsapPositionFromDrag from its actual module
The function was split out into gsapDragPositionCommit.ts in #1605, but the
test kept importing it from ./gsapDragCommit, which no longer exports it —
yielding 'is not a function' at runtime. Import from the correct module.
Inherited main breakage (same fix as #1631); fixes the Test CI check on this
branch independently of merge order.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(hyperframes): refine router skill metadata tags
Update the entry router's metadata tags (video / animation / router focus);
oxfmt collapses the now-shorter metadata to a single line.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(skills): tighten caption comment-strip + document audio --only merge
Review follow-ups (#1635):
- captions.mjs (x3): the HTML-comment strip used a single global replace, which
CodeQL flags as incomplete multi-character sanitization (a nested/partial pair
can re-form a marker the single pass misses). Strip in a fixpoint loop instead.
Input is preset-library content, not user-controlled, so this is lint-
cleanliness, not XSS defense.
- audio.mjs (x3): document that fetch-sfx (--only sfx) MERGES into the neutral
audio_engine_meta.json sidecar — the engine reads prev and recomputes only the
sfx section, so voices/bgm from the generate pass are preserved (review Q).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(skills): remove existsSync->write TOCTOU in workflow scripts
Clears the 9 js/file-system-race CodeQL alerts (captions/audio/transitions x3).
Each was an existsSync precheck followed by a later write of the same path:
- captions.mjs: caption-overrides shim -> atomic writeFileSync({ flag: 'wx' }).
- audio.mjs (sync-durations) + transitions.mjs (inject): drop the existsSync
precheck and read directly, surfacing the same friendly error from a try/catch
on readFileSync — no check->write gap.
Behavior is unchanged (same error messages); these are local single-process
deterministic scripts so the race was never a real risk, but this clears the gate.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(skills): paint root composition ground color in assemble-index
Per-frame roots carry data-start/data-duration and get clip-gated against the
global timeline at render, so only the first frame's window overlaps global 0 —
a frame's own full-bleed background can't serve as the video ground, and every
frame after the first renders on the bare body color (black). Paint the ground
on the always-present root composition using the project's frame.md canvas color
(the same role the caption skin maps to --cap-canvas); fall back to the body
letterbox color when frame.md is absent or has no resolvable ground.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(hyperframes): drop router-tag edit (moved to the foundation PR)
The entry SKILL.md is rewritten wholesale by the frame-presets/media foundation
PR (#1632); editing it here too guaranteed a merge conflict. Restore this file
to main and let the router-tag tweak live with the rewrite in #1632, so the two
PRs no longer both touch it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>