Files
hyperframes/skills/website-to-hyperframes/references/step-5-vo.md
T
Ular KimsanovandClaude Opus 4.7 a78ec5ec76 fix(skill): all file paths use capture/ subfolder prefix (#345)
* fix(skill): all file paths use capture/ subfolder prefix

Step-1 was updated to capture into <project-dir>/capture/ but all
file-read instructions across steps 1-6 still used bare paths
(extracted/tokens.json, assets/hero.png). Agents following the skill
looked for files in wrong locations, producing broken compositions
with 404'd images and fonts.

- step-1: all 12+ file-read paths prefixed with capture/
- step-2: tokens.json and asset-descriptions.md paths fixed
- step-3: visible-text.txt path fixed
- step-4: asset-descriptions.md, asset examples, directory tree expanded
  to show capture/ children (screenshots/, assets/, extracted/)
- step-5: minor cleanup
- step-6: asset-descriptions.md path fixed, composition examples already
  correct (../capture/assets/)
- techniques.md: Lottie, video, font paths all prefixed

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(skill): address Copilot review — font comment path, add videos/ to tree

- techniques.md: comment says ../capture/assets/fonts/ (matches composition context)
- step-4 file tree: add videos/ under capture/assets/

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-20 05:01:25 +02:00

2.4 KiB

Step 5: Generate VO + Map Timing

Audition voices

Never use the first voice you find. Audition 2-3 voices with the first sentence of SCRIPT.md:

  • Kokoro (try first — free, no API key) — npx hyperframes tts SCRIPT.md --voice af_nova --output narration.wav. Runs locally on CPU. Requires Python 3.10+ (macOS system Python 3.9 won't work — if it fails with an onnxruntime error, move to the next option).
  • ElevenLabs (best voice quality, widest selection) — mcp__elevenlabs__search_voices to browse, mcp__elevenlabs__text_to_speech to generate. Does not return timestamps — transcribe separately after.
  • HeyGen TTS (returns word timestamps automatically — saves a transcribe step) — mcp__claude_ai_HeyGen__text_to_speech. Use when you want timestamps without a separate transcription pass.

Pick the voice that sounds most natural and conversational. Listen for pacing — does it breathe between sentences? Does it sound like a person or a robot?

Generate full narration

Generate the full script as narration.wav (or .mp3) in the project directory.

Also save the exact spoken text — with pronunciation substitutions applied (e.g., APIA P I, $2Ttwo trillion) — as narration.txt in the same directory. This is the string passed to TTS, distinct from SCRIPT.md which is the human-readable creative doc. Having narration.txt makes it trivial to regenerate the audio later with a different voice without re-deriving the substitutions. Name it exactly narration.txt.

Transcribe for word-level timestamps

npx hyperframes transcribe narration.wav

Produces transcript.json with [{ text, start, end }] for every word. These timestamps are the source of truth for all beat durations.

Map timestamps to beats

Go through STORYBOARD.md beat by beat. For each beat:

  1. Find the first word of that beat's VO cue in transcript.json
  2. Find the last word of that beat's VO cue
  3. Set beat.start = firstWord.start, beat.end = lastWord.end
  4. Add 0.3-0.5s padding at the end for visual breathing room

Update STORYBOARD.md with real durations. Replace estimated times (e.g., "0:00-0:05") with actual timestamps (e.g., "0.00-3.21s").

Beat boundaries land on word onsets — hard cuts to the VO.

Update index.html

Update each scene slot's data-start and data-duration to match the real beat timings from the transcript. Also update the total composition duration and audio element duration.