Files
hyperframes/skills/hyperframes-media/references/tts.md
T
WaterrrForeverandClaude Opus 4.8 05af482f22 feat(skills): product-launch-video skill + consolidate motion knowledge into hyperframes-animation (#1745)
* feat(skills): product-launch-video + consolidate motion knowledge into hyperframes-animation

- Add the product-launch-video skill: shot-sequence architecture where each
  visual frame is a time-coded shot sequence picked from a blueprint menu and
  paced to the voiceover (anti-PowerPoint). Includes the frame-worker sub-agent,
  story/visual/motion-design references, and audio/captions/transitions/
  stage-assets/assemble-index scripts.
- Consolidate motion knowledge in hyperframes-animation as the single source of
  truth: promote the updated atomic rules (31 -> 36) and rename product-launch-
  video's archetypes into hyperframes-animation blueprints (13 -> 15, replacing
  the old set). product-launch-video, faceless-explainer, and pr-to-video now
  reference them via ../hyperframes-animation/{rules-index,blueprints-index}.md
  and the rules/blueprints dirs. Fixes the discrete-text-sequence broken links;
  blueprints no longer ship per-id runnable examples, so example references in
  the consumers were dropped.
- Default HeyGen TTS voice to Marcia (deterministic; was the API's first English
  voice, which drifts on catalog re-sort). Override with --voice.
- assemble-index pre-assembly frame guards: auto-repair a sub-comp root missing
  canvas dims; hard-fail on <video>/<audio> inside a sub-comp; hard-fail on a
  timed non-root element missing class="clip" or overlapping same-track clips.
- Lint/CLI: lint media inside sub-compositions as an error; stop false-positive
  caption layout/lint findings; contrast/layout-audit skip elements hidden by an
  invisible ancestor.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(skills): clear CodeQL alerts in assemble-index.mjs

- script/style blanking regex now matches closing tags with trailing
  whitespace (</script >, </style >) — js/bad-tag-filter (high).
- drop the existsSync precheck before reading/repairing a frame file; read
  directly and handle ENOENT, removing the check->write TOCTOU window —
  js/file-system-race (high).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-27 02:51:35 +08:00

7.4 KiB

Text To Speech

npx hyperframes tts auto-detects a provider from env vars; explicit override via --provider.

Run the Preflight first — no credential is not a green light to silently use the local voice. Before generating a voiceover, complete the sign-in Preflight (see ../SKILL.md → Preflight): run npx hyperframes auth status, recommend signing in, and STOP for the user's choice (sign in for HeyGen voices, or continue offline with local Kokoro). This applies to a one-off "generate a voiceover" request just as much as inside a full workflow.

Provider chain

Order Provider Env trigger Voice IDs Word timestamps Audio format
1 HeyGen (Starfish) $HEYGEN_API_KEY / ~/.heygen/credentials UUIDs from GET /v3/voices?engine=starfish Yes (word_timestamps[] in response) mp3 → wav via ffmpeg
2 ElevenLabs $ELEVENLABS_API_KEY UUIDs from elevenlabs.io dashboard No mp3 → wav via ffmpeg
3 Kokoro-82M always (local fallback) am_michael, af_heart, … (54 voices) No wav direct
# Auto-detect (HeyGen if key set, else ElevenLabs, else Kokoro)
npx hyperframes tts "Welcome to HyperFrames" -o narration.wav

# Pin the provider explicitly
npx hyperframes tts "Hello" --provider kokoro
npx hyperframes tts "Hello" --provider heygen --voice <heygen-uuid>
npx hyperframes tts "Hello" --provider elevenlabs --voice 21m00Tcm4TlvDq8ikWAM

# HeyGen path: capture word timestamps in one call (skips a Whisper pass)
npx hyperframes tts "Hi there" --words narration.words.json

Self-contained HeyGen (no CLI) — scripts/heygen-tts.mjs

The published hyperframes tts CLI synthesizes locally with Kokoro only. When you want HeyGen specifically — best quality plus word timestamps in one call — use the skill's bundled script, which calls the HeyGen v3 REST API directly and needs no CLI provider plumbing:

The script resolves a HeyGen credential the same way the CLI does — first source wins: $HEYGEN_API_KEY$HYPERFRAMES_API_KEY → a project .env (auto-loaded, walks up ≤5 dirs) → ~/.heygen/credentials (shared with heygen-cli; $HEYGEN_CONFIG_DIR overrides the dir). An OAuth login is sent as Authorization: Bearer; an API key as X-Api-Key. If the only credential is an expired OAuth token it stops with a hint to run npx hyperframes auth refresh.

# Only needed if you haven't run `npx hyperframes auth login`:
export HEYGEN_API_KEY=...   # or put it in a project .env

# Synthesize + capture word timestamps in one call (skips a Whisper pass)
node skills/hyperframes-media/scripts/heygen-tts.mjs \
  "Welcome to HyperFrames." -o narration.wav --words narration.words.json

node skills/hyperframes-media/scripts/heygen-tts.mjs ./script.txt -o narration.wav
node skills/hyperframes-media/scripts/heygen-tts.mjs --list   # public starfish voices
  • Voice: --voice <id> must be a starfish voice_id (--list, or GET /v3/voices?engine=starfish). v2-catalog ids are rejected with HTTP 400. Omit --voice (English) and it defaults to Marcia (05f19352e8f74b0392a8f411eba40de1, a fixed default so the choice is deterministic). Non-English with no --voice falls back to the first matching catalog voice.
  • Output: .wav → transcoded to 44.1k mono via ffmpeg; .mp3 → raw bytes (no ffmpeg needed).
  • Words: --words <path> writes the flat [{id,text,start,end}] shape below, drop-in for the captions pipeline. HeyGen's <start>/<end> boundary sentinels are filtered out and ids are re-contiguous.
  • Non-English: --lang <code> (anything but en) is sent as the request language.

When to use which provider

Goal Use
Best voice quality + word timestamps in one call HeyGen
Drop-in cloud TTS, big voice catalog ElevenLabs
Offline, no API key, fast iteration Kokoro
Non-English multilingual with deterministic phonemization Kokoro (ef_dora, jf_alpha, zf_xiaobei, …)

ffmpeg requirement

HeyGen + ElevenLabs return mp3. The CLI transcodes to wav when --output ends in .wav (the default and what downstream ffprobe + Whisper expect). If you'd rather skip the transcode, pass -o file.mp3. Without ffmpeg on PATH, .wav output from the cloud providers fails — install ffmpeg or use .mp3.

Voice selection (Kokoro)

Default af_heart. Curated picks:

Content type Voice
Product demo af_heart, af_nova
Tutorial / how-to am_adam, bf_emma
Marketing / promo af_sky, am_michael
Documentation bf_emma, bm_george
Casual / social af_heart, af_sky

Run npx hyperframes tts --list for the bundled set.

Multilingual (Kokoro voice prefix → language)

The first letter of a Kokoro voice ID picks the phonemizer language; --lang overrides auto-detection.

Prefix Language
a American English
b British English
e Spanish
f French
h Hindi
i Italian
j Japanese
p Brazilian Portuguese
z Mandarin
npx hyperframes tts "La reunión empieza a las nueve" --voice ef_dora --provider kokoro
npx hyperframes tts "Today is a nice day" --voice af_heart --provider kokoro

Valid --lang codes (only needed to override the voice's auto-detected language): en-us, en-gb, es, fr-fr, hi, it, pt-br, ja, zh.

Non-English phonemization requires espeak-ng system-wide (brew install espeak-ng / apt-get install espeak-ng).

Speed

  • 0.7-0.8 — tutorial, complex content, accessibility
  • 1.0 — natural pace (default)
  • 1.1-1.2 — intros, transitions, upbeat content
  • 1.5+ — rarely appropriate, test carefully

Honored by Kokoro + HeyGen; ElevenLabs ignores --speed (use voice settings on their dashboard).

Long scripts

Past a few paragraphs, write the text to a .txt file and pass the path. Inputs over ~5 minutes of speech may benefit from splitting into segments.

HeyGen word-timestamp shape

When --words <path> is passed to a HeyGen call, the file is written in the same flat shape transcribe produces — drop-in compatible with the captions pipeline:

[
  { "id": "w0", "text": "Hi", "start": 0.0, "end": 0.21 },
  { "id": "w1", "text": "there", "start": 0.22, "end": 0.55 }
]

For ElevenLabs / Kokoro, run npx hyperframes transcribe narration.wav --model small.en to get the same shape.