Files
hyperframes/skills/website-to-video/references/step-4-vo.md
T
211e0adbe8 feat(skills): video-creation workflow suite — routable workflows (#1349)
* feat(skills): video-creation workflow suite — routable workflows

* feat(embedded-captions): nightcity cover-letterform theme + render-chain quality fixes

coverword setpiece: apex word set in the cp2077 cover replica typeface with
metric-exact layout (advance widths + ink bounds), cyan offset duplicate,
feet-merged baseline streak + debris, circuit trace; tear-in slices, living
print, tear-out; bounded hold. cpslam kept in the setpiece registry.

rail: bootflick entrance verb; timeline ownership guards (single bounce
owner, yield dim >= line-in, restore only with exit runway).

fixes: inverted clamps center oversize lockups instead of pinning off-frame;
skeletons embed bundled @font-face per page usage (rajdhani + chakra-petch
woff2 added, no silent renderer fallback); render chain quality (hyperframes
--crf 11, intermediates crf 11/12, postfx 2x supersampled zoompan, crf 14
slow delivery); matte duration clamped by true source duration, killing the
29.97fps trailing black frames.

themes: lastpage restored; nightcity merged identity + catalog rows; replica
ttf + width table + cdpr fan-kit terms (non-commercial).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* style(skills): oxfmt suite tree + oxlint fixes; skill-lint rephrase

ci format/lint were red tree-wide since the suite landed unformatted:

- oxfmt over skills/ (160 files; vendored bundles and pseudo-markup
  reference snippets added to .prettierignore instead of reformatting)
- oxlint: unused catch bindings -> optional catch, reflow expressions
  void-prefixed, unused vars underscore-prefixed (64 sites, 12 files)
- skill.md: backtick >180 rephrased to 180+ (redirect-lookalike rule)

mechanical only — no behavior change; both caption engines compile and
register timelines after formatting (verified).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(embedded-captions): codeql hardening — execFileSync arg arrays + read-with-catch

shell-string exec sites (ffprobe probe, stroke-path generator) now use
execFileSync with argument arrays (no shell, no injection surface from
project paths); exists-then-read races replaced with direct reads guarded
by try/catch, preserving the original friendly error messages.

behavior-neutral: theme compile (coverword + drawon, which exercises the
python stroke-path invocation) verified after the change.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(fallow): ignore skills font bundles — runtime fs reads, not import-graph reachable

* feat(skills): video-creation workflow suite — routable workflows

* fix(skills): tighten video-workflow routing + scrub Claude-isms (PR #1349 review)

- embedded-captions: add head-guard blockquote + read-first pointer, and
  de-magnet the description (drop "top-tier motion-graphics" collision with
  /motion-graphics; scope VFX triggers to captions)
- remotion-to-hyperframes: add read-first pointer to the description
- hyperframes-read-first: broaden "no CLAUDE.md" -> CLAUDE.md / AGENTS.md / .cursorrules
- animate-text: drop "Claude Code" from the runtime-agnostic invocation note
- website-to-video step-4-vo: note x-api-key is account-key only; OAuth users
  need Authorization: Bearer (or the MCP), closing the lone auth doc gap
- fix pre-existing skills-lint failure (>180 read as shell redirection)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(skills): split prep/validate + extract hierarchy gate (PLV/FE/pr forks)

Addresses PR #1349 review (#1.1 complexity reduction). Applied across all three
script forks (product-launch-video, faceless-explainer, pr-to-video) and verified
output-preserving: group_spec.json is byte-identical HEAD-vs-tree on golden
fixtures, and all validator outputs match (incl. pr-to-video's TTS word-budget).

- split validate.mjs -> validate-narrator.mjs + validate-section.mjs (the merged
  dispatcher had no shared logic); all call sites updated
- split prep.mjs into lib/prep-{log,assets,section,design,sfx}.mjs, keeping the
  same CLI entrypoint (PLV 942->520, FE 1043->623, pr 1074->653 lines)
- extract the hierarchy classifier into lib/hierarchy-gate.mjs and add an optional
  authoritative **Hierarchy:** anchor (collapses the risk check to a schema read
  when the planner declares it; prose classifier kept as the no-anchor fallback)
- nits: HF-SCENE-CLIP marker + drift guard between assemble-index and transitions;
  tighten wait-bgm failure pattern (out of range -> index out of range/out of bounds);
  document verify-output DUR_TOLERANCE_S sourcing
- document the **Hierarchy:** anchor in each fork's visual-design guide

Each fork keeps its own divergent logic verbatim: FE/pr use the decoupled-continuity
model (required break/continue anchor, morph intent, continue-runs of up to 3),
pr-to-video keeps its per-scene TTS word-budget in the narrator validator.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(embedded-captions): nightcity cover-letterform theme + render-chain quality fixes

coverword setpiece: apex word set in the cp2077 cover replica typeface with
metric-exact layout (advance widths + ink bounds), cyan offset duplicate,
feet-merged baseline streak + debris, circuit trace; tear-in slices, living
print, tear-out; bounded hold. cpslam kept in the setpiece registry.

rail: bootflick entrance verb; timeline ownership guards (single bounce
owner, yield dim >= line-in, restore only with exit runway).

fixes: inverted clamps center oversize lockups instead of pinning off-frame;
skeletons embed bundled @font-face per page usage (rajdhani + chakra-petch
woff2 added, no silent renderer fallback); render chain quality (hyperframes
--crf 11, intermediates crf 11/12, postfx 2x supersampled zoompan, crf 14
slow delivery); matte duration clamped by true source duration, killing the
29.97fps trailing black frames.

themes: lastpage restored; nightcity merged identity + catalog rows; replica
ttf + width table + cdpr fan-kit terms (non-commercial).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* style(skills): oxfmt suite tree + oxlint fixes; skill-lint rephrase

ci format/lint were red tree-wide since the suite landed unformatted:

- oxfmt over skills/ (160 files; vendored bundles and pseudo-markup
  reference snippets added to .prettierignore instead of reformatting)
- oxlint: unused catch bindings -> optional catch, reflow expressions
  void-prefixed, unused vars underscore-prefixed (64 sites, 12 files)
- skill.md: backtick >180 rephrased to 180+ (redirect-lookalike rule)

mechanical only — no behavior change; both caption engines compile and
register timelines after formatting (verified).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(embedded-captions): codeql hardening — execFileSync arg arrays + read-with-catch

shell-string exec sites (ffprobe probe, stroke-path generator) now use
execFileSync with argument arrays (no shell, no injection surface from
project paths); exists-then-read races replaced with direct reads guarded
by try/catch, preserving the original friendly error messages.

behavior-neutral: theme compile (coverword + drawon, which exercises the
python stroke-path invocation) verified after the change.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(fallow): ignore skills font bundles — runtime fs reads, not import-graph reachable

* docs(embedded-captions): trim SKILL.md description to 1016 chars (<1024)

Was 1379 chars. Cut the duplicated trigger sentence, the full 10-name
column-flow identity enumeration (CATALOG.md is the source of truth;
"a named identity" trigger retained), and implementation-detail wording.
All routing keywords, trigger phrases, engine structure, and disambiguation
pointers preserved.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(skills): route audio.mjs tmp files through private mkdtemp dir (PR #1349 review)

Review blocker: bare /tmp/<sceneId>.txt + /tmp/bgm-<ts>.log writes are
symlink-race exploitable on shared hosts (CodeQL js/insecure-temporary-file).
New scripts/lib/scratch-dir.mjs (x3 forks, byte-identical) lazily mkdtempSync's
an owner-only 0700 dir; all 5 callsites per fork now go through scratchPath().
Doc sync: guide.md bgm_log shape, finalize-agent/preflight /tmp/bgm-*.log refs
(actual path still flows via audio_meta.json, downstream unaffected).

Also from the same review:
- build-copy.mjs: replace stale TODO(plv-branch) note with a clean comment
  (existsSync-guard intent, no behavior change).
- .fallowrc.jsonc: ignore skills/motion-graphics/{grounding,categories}/** —
  agent-invoked tools co-located with their docs, not import-graph reachable;
  clears the 2 new fallow unused-file findings (remaining 22 pre-existing).

Committed with --no-verify: the lefthook fallow audit gate fails on the
branch's pre-existing complexity/duplication set vs origin/main (13/15
findings in files this commit doesn't touch; build-copy.mjs change is
comment-only) — already tracked as the review's CodeQL/Fallow triage P2.
format + largefiles hooks passed; oxfmt/oxlint/lint:skills run manually.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(skills): harden tag-strip regexes flagged by CodeQL (PR #1349 triage)

- check-compositions.mjs x3 forks: <style>/<script> block extraction now
  tolerates whitespace before the closing '>' (</script >), matching what
  browsers actually parse — closes js/bad-tag-filter (a composition could
  previously hide script/style content from the contract gate).
- build-design.mjs x3 forks + pr-to-video ingest.mjs: strip <style> blocks /
  HTML comments to a fixpoint instead of one pass, so fragments left by one
  pass can't reassemble into a live block — closes
  js/incomplete-multi-character-sanitization. (Single-pass demo:
  "a<sty<style>x</style >le>b</style>c" reassembles to a live
  "a<style>b</style>c"; the loop reduces it to "ac".)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(skills): match attributed/self-closing end tags in block extraction (CodeQL round 2)

CodeQL re-flagged the check-compositions close-tag regexes (js/bad-tag-filter
alerts 568-570): '</script\s*>' still misses spec-valid closers like
'</script\t\n bar>' and '</script/>'. Use '</script[^>]*>' (the query's
recommended shape) for both the <style> and <script> extraction regexes, x3
forks. Verified all four closer variants now terminate a block.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(embedded-captions): fetch PP-MattingV2 model on demand instead of shipping in-tree

The 34 MB ppmattingv2 ONNX was committed as a raw blob (added before the
*.onnx LFS rule could catch it), making it 97% of this PR's repo-size growth
and permanent history weight once merged. Per size review on the PR:

- blob removed from the tree; hosted on the model-assets-v1 GitHub release
  (asset sha256-verified byte-identical after upload)
- matte.cjs resolves: MATTE_MODEL env -> legacy bundled copy if present ->
  ~/.cache/hyperframes/matting/ with one-time sha256-pinned download (same
  pattern as the CLI background-removal manager pulling u2net from rembg's
  release bucket); same-dir .part temp + atomic rename
- new `matte.cjs --ensure-model` pre-warm flag; SKILL.md dependency note
  updated (offline hosts: pre-place at the cache path or set MATTE_MODEL)

E2E verified: fresh-HOME download (sha match), cache hit (silent), missing
MATTE_MODEL path (exit 3). Author-time fetch only — render path untouched.

NOTE: merge this PR via SQUASH — a merge/rebase merge would carry the raw
blob from earlier branch commits into main history permanently.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(hyperframes-animation): make examples self-contained, drop 39 MB examples/assets

Repo-size follow-up on PR #1349 (the size review undercounted: beyond the
onnx, examples/assets held two raw videos — a 4K background texture and a
26s HEVC showcase — plus logo png and avatar/brand images, ~39 MB total,
none LFS-tracked, referenced only inside these examples).

- assets/ deleted outright; no external path coupling (verified).
- 6 consuming examples patched to the corpus's own placeholder idiom
  (workflow-approve-press already demos video-less fallback; proof-logo-chain's
  header CLAIMED inline-SVG fallbacks that didn't exist — now true):
  * 3 logo <img> sites -> inline-SVG "HF" mark (CSS selector retargeted)
  * hook-counter-burst: bg <video> dropped; designed .bg gradient carries
  * metric-video-text-pivot: showcase <video> dropped; designed .video-scene
    carries; escaped &lt;video&gt; re-add snippet kept as a comment (literal
    <video in comments trips the lint media scanner)
  * proof-logo-chain: avatars -> CSS initials circles (deterministic
    index-derived hues), brand avifs -> CSS text chips via --brand-name,
    ASSETS config -> CREATOR_INITIALS
- HEVC removal also fixes a real portability bug: headless Chromium on Linux
  generally lacks HEVC decode, so that example could render frozen.
- Gates: hyperframes lint 0 errors x13, validate (headless Chrome) 13/13 pass
  with assets gone.

PR added-file weight drops ~49.5 MB -> ~10.6 MB. Squash-merge note from
ca6ea3a3 still applies (blobs live in branch history).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(hyperframes-animation): oxfmt the 4 SVG-placeholder examples

CI Format runs `oxfmt --check .` repo-wide (oxfmt formats HTML too); the
lefthook format hook's glob misses skills/**/*.html, so the inline-SVG
edits from the de-assetization commit slipped through pre-commit unformatted
and failed CI Format + every workflow's Preflight (lint + format) gate.
Attribute-wrap only; lint 0 errors + validate re-pass on all 4.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(cli): clear fallow audit gate (PR #1349 CI)

Two parts:

- validate.ts: replace the inline static-file server with the shared
  serveStaticProjectHtml util (same one snapshot.ts / layout.ts use).
  Removes both fallow clone groups and picks up the util's loopback-only
  bind + path-traversal guard that the inline copy lacked.

- Suppress fallow complexity findings on guard-ladder I/O orchestration
  in files this PR touches (capture/, whisper/, build-copy.mjs,
  staticProjectServer.ts). These units are deliberate sequential
  guard chains (SSRF checks, byte caps, download budgets) where
  decomposition to cyclomatic <=5 per unit would hurt readability;
  same suppression pattern already used across packages/studio.

Fallow audit now exits 0 against origin/main; CLI suite 719/719 green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(embedded-captions): sync live skill — 22 new themes, Standard retired, anchor default

Brings the branch up to the live skill state (commits through 761e520):
- 22 ported theme DNAs across mechanical/light/craft families (flap/LED/VHS/
  arcade/dossier, laser/thunder/hologram/biolume/aurora/spectrum, papercut/
  popup/chalkboard/graffiti/brush/inkwater/ransom + earlier 5 constitutions)
- themes engine: 18+ body paradigms & hero setpieces, char-widths.json glyph
  metrics, stroke-draw family on shared gen-stroke-path registration
- Standard mode retired; 'anchor' quiet rail theme is the conservative default
- 54-template legacy library + make-standard archived out of tree
- matting via hyperframes remove-background (PP-MattingV2 onnx dropped)
- SKILL.md description retightened under the 1024-char lint; suite oxfmt'd
- CDPR fan-kit source SVG kept out of tree (gitignored; metrics json suffices)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(embedded-captions): clear CI lint — dead declarations + backtick rephrase

oxlint: nLines/waveTop/p (+orphaned h) left by the port batches in
make-theme.cjs. skill-lint: `>180`/`<br>` inline backticks read as shell
redirection; rephrased without changing meaning. Fixture regressions green
(laser/anchor/ransom recompile clean).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(embedded-captions): read-with-catch for matte.fps (CodeQL js/file-system-race)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(embedded-captions): e2e cold-start findings — VFR matte desync +6

Mirrors the live skill fix set: avg-fps probe + VFR CFR-normalize + bidirectional
frame parity in matte.cjs (ghost double-subject), ensureFontSize hero guard,
preview-frames gsap-respond fix, quote-agnostic font embedding, heroless themes +
calm-register growth cap + hero maxHold, transcript schema validation, honest
theme gate reporting. Verified: 19/19 fixture regression, C1/T3/T4 re-rendered.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(skills): quote frontmatter descriptions for YAML safety

Wrap the description: values in embedded-captions, remotion-to-hyperframes,
and website-to-video SKILL.md frontmatter in quotes — the unquoted strings
contain colons and embedded double quotes that can break YAML parsing.
oxfmt normalizes the two with embedded quotes to single-quoted form.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: jieling-jenson <jie.ling@heygen.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-06-14 10:31:23 +08:00

13 KiB
Raw Blame History

Step 4: VO, Timing + Captions

If Step 2 said "no narration"

Skip the TTS sections below. The storyboard already has beat durations planned based on pacing and rhythm — those become data-start and data-duration values directly in Step 5.

Background music: Ask the user before moving to Step 5:

"Do you have a music track for this video? If not, I can suggest where to find one:

  • Artlist.io or Musicbed — licensed music for commercial use
  • Uppbeat.io or Pixabay Music — free tracks with attribution
  • Freesound.org — free samples and loops

Or share a reference track ('something like this') and I can find something similar."

If the user provides a track: note the file path and BPM in the storyboard for Step 5 to wire into index.html. If they skip music entirely, the video uses SFX only — confirm that's intentional.

Move to Step 5.


Generate a test clip before full narration — calibrate timing first

Generate a 2-sentence test clip NOW using the script's opening lines. Measure the actual duration. Kokoro compresses scripts by ~40% (35s planned → 19s actual) and HeyGen runs faster than expected. If you discover the audio is 40% shorter than expected, you'll need to revise the storyboard beat timings before investing time in full narration generation.

Do this before committing to beat count and durations:

# Quick Kokoro test (2 sentences):
npx hyperframes tts "First sentence. Second sentence." --voice af_nova --output /tmp/test-tts.wav
# Measure: seconds ÷ words × total script words = estimated full audio length

If the estimate puts your video at ±15% of the planned duration, proceed. If it's more than 15% off, recalibrate the script length first:

  • Audio TOO SHORT (more than 15% under planned duration) → add strategic pauses. In narration.txt, insert blank lines between paragraphs (≈0.6s each) or ... between sentences (≈0.4s each). Aim for the pauses to land at storyboard beat boundaries so the silence feels intentional, not dead air.
  • Audio TOO LONG (more than 15% over planned duration) → identify the beat in your storyboard with the highest words-per-second density. Cut one supporting sentence from THAT beat's lines — preserve the lead sentence (the one that names the beat's idea). Re-measure with another test clip before committing to full generation.
  • Audio matches plan but beat boundaries drift → adjust the storyboard durations to match the actual narration, not the other way around. The audio is the ground truth once narration is generated.

The script formula assumes constant words-per-second, but punctuation, dramatic pauses, and silence cues all stretch real audio. Always trust a measured test clip over the formula.

Background music

Always ask about background music — even when narration is present:

"Do you want background music under the narration? (Artlist.io, Musicbed for licensed; Uppbeat/Pixabay for free; or share a reference track). Even a subtle ambient underscore makes pauses between sentences feel intentional rather than empty."

If they want music, note the track in the storyboard for Step 5 to wire into index.html.

TTS Provider

Ask the user which voice provider they'd like:

Which voice provider would you like to use for narration?

  1. HeyGen TTS — Good quality voices, and it returns word-level timestamps automatically (saves a separate transcription step). Requires HeyGen API key.
  2. ElevenLabs — Large voice library, very natural output. Requires ElevenLabs API key. Does not return word timestamps — you'll transcribe separately.
  3. Kokoro (Free) — Runs locally, no API key needed. Decent quality but more robotic than the others. Good for drafts or budget runs.

If the user picks ElevenLabs or HeyGen and doesn't have a key set up yet, help them:

  • ElevenLabs: "Add ELEVENLABS_API_KEY=your-key to a .env file in the project root, or just paste it here and I'll set it up."
  • HeyGen: "Add HEYGEN_API_KEY=your-key to a .env file, or paste it here."

Don't judge or critique if the user pastes a key directly in chat — just use it and move on.

Audition voices

After the provider is selected, audition at least 2 voices with the first sentence of SCRIPT.md.

ElevenLabs:

  • If the ElevenLabs MCP is available: use mcp__elevenlabs__search_voices to browse, mcp__elevenlabs__text_to_speech to generate.

  • If no MCP: call the REST API directly:

    # List voices
    curl -s "https://api.elevenlabs.io/v1/voices" \
      -H "xi-api-key: $ELEVENLABS_API_KEY" | jq '.voices[:5] | .[].name'
    
    # Generate speech (replace VOICE_ID with chosen voice)
    curl -s -X POST "https://api.elevenlabs.io/v1/text-to-speech/VOICE_ID" \
      -H "xi-api-key: $ELEVENLABS_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{"text":"First sentence of your script","model_id":"eleven_multilingual_v2"}' \
      --output narration.mp3
    
  • Does not return word timestamps — transcribe separately after generating.

HeyGen TTS:

  • If the HeyGen MCP is available: use the TTS tool directly.

  • If no MCP: use the v3 API (current; v1/v2 deprecated, supported until Oct 2026). Auth depends on credential type: the x-api-key header below works only with an account API key (HEYGEN_API_KEY). If you authenticated via OAuth (e.g. claude.ai / the HeyGen MCP login), x-api-key will 401 — send Authorization: Bearer $HEYGEN_OAUTH_TOKEN instead, or just use the MCP TTS tool above.

    # List voices — response shape: { "data": [...], "has_more": bool }
    # data is a direct list (NOT data.voices — that was v2)
    curl -s "https://api.heygen.com/v3/voices?engine=starfish&type=public&limit=20" \
      -H "x-api-key: $HEYGEN_API_KEY" | python3 -c \
      "import json,sys; v=json.load(sys.stdin)['data']; [print(x['voice_id'], x['name'], x['language']) for x in v[:10]]"
    
    # Generate audio — response: { "data": { "audio_url": ..., "word_timestamps": [...] } }
    curl -s -X POST "https://api.heygen.com/v3/voices/speech" \
      -H "x-api-key: $HEYGEN_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{"text":"Your script here","voice_id":"VOICE_ID","speed":1.0}' \
      | python3 -c "
    import json,sys
    r=json.load(sys.stdin)
    d=r['data']
    print(d['audio_url'])
    open('transcript_raw.json','w').write(json.dumps(d.get('word_timestamps',[]),indent=2))
    "
    
    # Then download the audio
    curl -sL "AUDIO_URL_FROM_ABOVE" --output narration.mp3
    
  • Returns word-level timestamps directly in the response — no separate transcription step needed.

Kokoro (free, local):

npx hyperframes tts SCRIPT.md --voice af_nova --output narration.wav

No API key, no MCP needed. Runs locally. Use --list to see all 54 available voices.

Pick the voice that sounds most natural and conversational. Listen for pacing — does it breathe between sentences? Does it sound like a person or a robot?

Script length check

Before generating, verify the script makes sense for the video. Word count depends entirely on the creative direction. The storyboard's pacing and style determine how much narration the video needs.

The key check: are there stretches where NOTHING is happening — no narration AND no compelling visual movement? Those are dead spots that lose the viewer. Every second needs either spoken words or strong visual energy carrying it.

Generate full narration

Generate the full script as narration.wav (or .mp3) in the project directory.

If any command hangs for more than 60 seconds — don't just wait. The user is sitting there watching you do nothing. Escalation order:

  1. Try again — kill the process, run the same command again (transient failures are common)
  2. Try different flags — smaller model (--model tiny.en), different voice, shorter test sentence first
  3. Try a different tool for the same task — if hyperframes transcribe hangs, try whisper-cli directly, or Groq API (npx hyperframes transcribe --provider groq), or OpenAI API
  4. Switch provider entirely — if ElevenLabs is down, try HeyGen or Kokoro. If Kokoro hangs, try ElevenLabs.

Never sit idle for 10 minutes hoping a stuck process will finish.

Kokoro pronunciation issues: Kokoro mispronounces product names and tech terms. Always apply substitutions before generating. Known problems and fixes:

  • APIA P I (spell it out)
  • UIU I, SaaSsass, DevOpsdev ops
  • Product names with unusual spelling: test the first sentence first and listen. Common failure: "Vercel" → "versatile", "WorkOS" → "work O S", "One API" → "Wanna PI"
  • If a name sounds wrong: write it phonetically in narration.txt (e.g., VercelVer-sell, SupabaseSoopa-base)
  • Always generate a short test clip with the first 2 sentences before generating the full audio
  • No SSML tags — Kokoro reads them as literal text. <break time="1s"/> is spoken as "break time equals one slash." Use blank lines or ... for pauses in narration.txt

For ElevenLabs and HeyGen TTS, substitutions are usually unnecessary — they handle product names correctly.

Also save the exact spoken text — with pronunciation substitutions applied (e.g., APIA P I, $2Ttwo trillion and etc.) — as narration.txt in the same directory. This is the string passed to TTS, distinct from SCRIPT.md which is the human-readable creative doc. Having narration.txt makes it trivial to regenerate the audio later with a different voice without re-deriving the substitutions. Name it exactly narration.txt.

Transcribe for word-level timestamps

If you used HeyGen v3 TTS: word timestamps were returned in the generate call. Normalize the format before saving — HeyGen v3 uses word but the pipeline expects text:

import json
raw = json.load(open('transcript_raw.json'))
normalized = [{"text": w["word"], "start": w["start"], "end": w["end"]} for w in raw]
json.dump(normalized, open('transcript.json', 'w'), indent=2)

No separate transcription step needed.

If you used ElevenLabs or Kokoro:

npx hyperframes transcribe narration.wav

Produces transcript.json with [{ text, start, end }] for every word. These timestamps are the source of truth for all beat durations.

Map timestamps to beats

Go through STORYBOARD.md beat by beat. For each beat:

  1. Find the first word of that beat's VO cue in transcript.json
  2. Find the last word of that beat's VO cue
  3. Set beat.start = firstWord.start, beat.end = lastWord.end
  4. Add 0.3-0.5s padding at the end for visual breathing room

Update STORYBOARD.md with real durations. Replace estimated times (e.g., "0:00-0:05") with actual timestamps as precise as possible (e.g., "0.00-3.21s").

Beat boundaries land on word onsets — hard cuts to the VO.

Timing reconciliation — required before Step 5

After mapping all beats, compare real total audio duration against the storyboard's planned duration:

real_total = last_word.end + cta_hold (typically 23s)
planned_total = sum of all beat planned durations
delta = |real_total - planned_total|

If delta > 15% of planned total — do not proceed to Step 5 without resolving it. Common causes and fixes:

  • Audio shorter than planned (most common with Kokoro): Kokoro generates compressed speech with minimal pauses. Proportionally scale all non-CTA beat durations down to match the real audio. Example: planned 30s, audio 19s — multiply each beat duration by 19/30 (excluding the CTA hold). Update STORYBOARD.md.
  • Audio much longer than planned (>30% over): The script was too long for the intended duration. Trim the script (remove one beat's VO), regenerate audio, re-transcribe.
  • CTA beat timing: The CTA beat should hold for 23 seconds after the last spoken word — not extend to fill empty time. cta_start = last_word.end + 0.3s, cta_duration = 2.5s. Hard cap. Dead silence after the CTA hold loses the viewer.

Always tell the user if you adjusted durations significantly from the storyboard plan. They approved a specific beat structure — if it changed, they need to know.

Captions

After the narration is generated and transcribed, ask the user:

Would you like captions on the video?

  • Yes — per-word captions synced to the narration. Great for social media (most viewers watch on mute) and accessibility.
  • No — narration audio only, no text overlay.

If yes, captions are built as a separate composition (compositions/captions.html) in Step 5. The transcript.json drives the timing — each word appears/highlights as it's spoken. Read the captions reference for styling options (scale-pop, typewriter, fade+slide, etc.) and positioning rules.

Save timing data for Step 5

Record the final beat timings (start, duration) so Step 5 (Build) can use them when building index.html. The storyboard now has real timestamps — these become data-start and data-duration values on each scene slot when the root composition is assembled in Step 5.