Files
hyperframes/.claude/skills/changelog-video/references/script-voice.md
T
James RussoandClaude Opus 4.7 807078c7cd docs(skills): make captions non-optional in changelog-video (#2729)
* docs(skills): make captions non-optional in changelog-video

The Jul 13-20 build shipped without captions because the pipeline had two soft
failure modes:

- Empty vo-words.json from the TTS provider was silently accepted, so
  align-captions.mjs had nothing to align. captions.json was never produced.
- Step 5 tolerated the master-skeleton's LINES array staying empty. The build
  lint-checked and validated clean, but the caption rail rendered no text.

This adds three hard gates on top of the existing scaffold:

- Step 4 adds a whisper forced-alignment fallback so a missing TTS timestamp
  payload no longer breaks the caption pipeline. Whisper only supplies timings;
  captions still use the DISPLAY layer from script-tokens.json.
- Step 5 flags an empty LINES array as a shipped bug and shows the exact
  IIFE-input shape it expects.
- Step 6 adds gate 5: sample 3-4 frames across the VO window and confirm visible
  caption text on each. If any spoken window renders no caption, the build is
  red.

Anti-patterns table gets two rows covering the empty-LINES and
missing-word-timings failure modes. .claude and .agents mirrors stay
byte-identical.

Follow-up to #2669 (pre-build gate); paired with the v5 rebuild that
retrofitted captions onto the shipped Jul 13-20 video.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* docs(skills): bump changelog caption size to 32px for phone-screen legibility

25px reads too small at 1080² when the video is viewed on a phone without
zoom. Bumped to 32px (font-size 25→32, top 1002→990, height 40→52, alpha
.92→.94) so captions carry across small-screen playback. Change is CSS-only
on the master-skeleton; existing per-scene chrome and layouts are untouched.

Retrofitted the Jul 13-20 v6 build with the same values (identical CSS
block) so the shipped video and the future default match. .claude and
.agents mirrors stay byte-identical.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* docs(skills): sync caption position refs to top: 990 across SKILL.md + build-spec

Miga caught: the master-skeleton bump to top: 990 / 32px / height: 52 was
not reflected in the SKILL.md prose or build-spec.md. All three files now
name the same values so the doc, gate, and code stay consistent. Also
clarifies build-spec.md's caption-rail line to point at SKILL.md step 5 as
the source of the LINES-populated requirement.

.claude and .agents mirrors byte-identical.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* docs(skills): sync script-voice.md caption rail to top: 990 / 32px

Third and final stale reference to top: 1002 / 25px — script-voice.md was
still stating the pre-bump values. Now matches master-skeleton.html,
build-spec.md, and SKILL.md. Also bumps the 'keep critical small text out'
guidance from ~80px to ~100px to reflect the slightly taller caption rail.

Full-directory sweep confirms zero remaining occurrences of 1002 or 25px
across .claude/ and .agents/ skill mirrors.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-21 20:38:46 -04:00

112 lines
4.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Script + Voice: the two-layer contract
The script is the single source of truth for BOTH the VO and the captions,
written as token lines. The VO reads the `spoken` layer; captions render the
`display` layer. This is a hard quality gate: a caption showing "jay-sawn" or
a VO saying "juh-son" ("JSON" read literally) are both build failures.
## Register (how it should sound)
- Conversational, not release-notes. "The big one this week —" beats
"Theme 1:". Contractions welcome. Second person allowed ("your clips").
- Informational, never salesy; no superlatives the changelog doesn't earn.
- One breath per beat: sentences ≤ ~14 words; let punctuation pace the read.
- Numbers with meaning stay ("fifteen releases"); commit hashes, PR numbers,
and version micro-detail are never spoken.
- Open with the week + the marquee, close with the digest pointer
("See everything at hyperframes dot heygen dot com").
- **Teach the simple command.** When a feature has a one-line invocation (a
slash command, a CLI one-liner), the script says it verbatim ("start your
prompt with /figma…") and the mock shows it being typed — the command is
the visible CAUSE of the result. Slash commands speak as "slash <name>",
caption as `/name`.
## Token-line format (`script-tokens.json`)
```json
{
"lines": [
{
"id": "l1",
"tokens": ["This", "week", "at", "HyperFrames,", { "display": "JSON", "spoken": "jay-sawn" }]
}
]
}
```
- A bare string = display and spoken identical.
- An object = the layers diverge. `display` keeps standard spelling AND the
punctuation captions should show; `spoken` is what the TTS reads.
- One line = one caption phrase (≤ ~40 chars of display text). Line grouping
is an authoring decision made here, not downstream.
- Build `vo-spoken.txt` by joining every token's spoken form with spaces,
lines joined into sentences/paragraphs as punctuated.
## Phonetics rules (ElevenLabs-style best practices, plain-text)
HeyGen TTS takes plain text (no SSML), so pronunciation is controlled by
spelling, hyphens, and spacing:
1. **Initialisms** (each letter said): space or hyphen the letters —
`CLI → "C L I"`, `CDP → "C D P"`, `API → "A P I"`.
2. **Acronyms said as words**: respell phonetically —
`JSON → "jay-sawn"`, `GSAP → "jee-sap"`.
3. **Mixed / pronounceable compounds**: hyphenated LOWERCASE phonetics, one
fluid run — `ffmpeg → "ff-mpeg"` (ear-tested; the TTS reads "ff" as a
fluid "eff-eff"), `WebM → "web em"`, `OAuth → "oh-auth"`. Never spaced
capitals here: the TTS reads spaced caps as isolated letter names with
hard stops ("F F em-peg" comes out "eff… eff… em-peg"). Reserve spaced
capitals for TRUE initialisms (CLI, API) where a deliberate
letter-by-letter read is the goal. When candidates are close, generate
A/B takes of the real sentence and let the user pick by ear.
4. **Versions/numbers**: expand — `v0.7.36 → "version zero point seven
point thirty-six"` (usually: don't speak versions at all),
`1080×1080 → "ten-eighty by ten-eighty"`.
5. **URLs**: `hyperframes.heygen.com → "hyperframes dot hey-jen dot com"`.
6. **Filenames/extensions**: `.mp4 → "dot em pee four"` — or rephrase so the
extension isn't spoken.
7. **Emphasis/pauses**: commas and em-dashes, never caps. Ellipses are
unreliable in TTS — use an em-dash.
The shared vocabulary lives in `references/lexicon.json`
(`display → spoken`). Consult it for EVERY technical term; if a term is
missing, ask the user for the pronunciation and add the entry — never guess,
never ship unheard. New entries: listen to that line in the generated VO
before accepting.
## Alignment (spoken timestamps → display captions)
`heygen-tts.mjs --words` returns word timestamps of the SPOKEN text.
`scripts/align-captions.mjs` walks the spoken stream against the token lines
(one display token may cover several spoken words — "C L I" is three) and
emits `captions.json`:
```json
{
"lines": [
{
"end": 3.1,
"w": [
["This", 0.22],
["week", 0.4],
["JSON", 1.1]
]
}
]
}
```
Each display word carries the start time of its FIRST spoken word; a line's
`end` = the next line's start (last line: last word end + 0.6). The aligner
warns `MISMATCH` when the heard word doesn't fuzzy-match the expected spoken
form — every warning must be resolved (fix the lexicon spelling or the
transcript) before the captions are trusted.
## Caption rail (rendering)
Per `captions-overlay`: a quiet OVERLAY, never a reserved band. One line,
bottom-center (top: 990px, height: 52px on 1080-square), TT Norms Pro 500 32px,
ink .94, soft dark text-shadow, words fading in (0.12s) on their timestamps,
phrase swaps as sets. Keep critical small text out of the bottom ~100px
center span; everything else may run under the rail.