mirror of
https://github.com/heygen-com/hyperframes.git
synced 2026-09-04 16:42:27 +00:00
* docs(skills): make captions non-optional in changelog-video The Jul 13-20 build shipped without captions because the pipeline had two soft failure modes: - Empty vo-words.json from the TTS provider was silently accepted, so align-captions.mjs had nothing to align. captions.json was never produced. - Step 5 tolerated the master-skeleton's LINES array staying empty. The build lint-checked and validated clean, but the caption rail rendered no text. This adds three hard gates on top of the existing scaffold: - Step 4 adds a whisper forced-alignment fallback so a missing TTS timestamp payload no longer breaks the caption pipeline. Whisper only supplies timings; captions still use the DISPLAY layer from script-tokens.json. - Step 5 flags an empty LINES array as a shipped bug and shows the exact IIFE-input shape it expects. - Step 6 adds gate 5: sample 3-4 frames across the VO window and confirm visible caption text on each. If any spoken window renders no caption, the build is red. Anti-patterns table gets two rows covering the empty-LINES and missing-word-timings failure modes. .claude and .agents mirrors stay byte-identical. Follow-up to #2669 (pre-build gate); paired with the v5 rebuild that retrofitted captions onto the shipped Jul 13-20 video. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * docs(skills): bump changelog caption size to 32px for phone-screen legibility 25px reads too small at 1080² when the video is viewed on a phone without zoom. Bumped to 32px (font-size 25→32, top 1002→990, height 40→52, alpha .92→.94) so captions carry across small-screen playback. Change is CSS-only on the master-skeleton; existing per-scene chrome and layouts are untouched. Retrofitted the Jul 13-20 v6 build with the same values (identical CSS block) so the shipped video and the future default match. .claude and .agents mirrors stay byte-identical. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * docs(skills): sync caption position refs to top: 990 across SKILL.md + build-spec Miga caught: the master-skeleton bump to top: 990 / 32px / height: 52 was not reflected in the SKILL.md prose or build-spec.md. All three files now name the same values so the doc, gate, and code stay consistent. Also clarifies build-spec.md's caption-rail line to point at SKILL.md step 5 as the source of the LINES-populated requirement. .claude and .agents mirrors byte-identical. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * docs(skills): sync script-voice.md caption rail to top: 990 / 32px Third and final stale reference to top: 1002 / 25px — script-voice.md was still stating the pre-bump values. Now matches master-skeleton.html, build-spec.md, and SKILL.md. Also bumps the 'keep critical small text out' guidance from ~80px to ~100px to reflect the slightly taller caption rail. Full-directory sweep confirms zero remaining occurrences of 1002 or 25px across .claude/ and .agents/ skill mirrors. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
112 lines
4.8 KiB
Markdown
112 lines
4.8 KiB
Markdown
# Script + Voice: the two-layer contract
|
||
|
||
The script is the single source of truth for BOTH the VO and the captions,
|
||
written as token lines. The VO reads the `spoken` layer; captions render the
|
||
`display` layer. This is a hard quality gate: a caption showing "jay-sawn" or
|
||
a VO saying "juh-son" ("JSON" read literally) are both build failures.
|
||
|
||
## Register (how it should sound)
|
||
|
||
- Conversational, not release-notes. "The big one this week —" beats
|
||
"Theme 1:". Contractions welcome. Second person allowed ("your clips").
|
||
- Informational, never salesy; no superlatives the changelog doesn't earn.
|
||
- One breath per beat: sentences ≤ ~14 words; let punctuation pace the read.
|
||
- Numbers with meaning stay ("fifteen releases"); commit hashes, PR numbers,
|
||
and version micro-detail are never spoken.
|
||
- Open with the week + the marquee, close with the digest pointer
|
||
("See everything at hyperframes dot heygen dot com").
|
||
- **Teach the simple command.** When a feature has a one-line invocation (a
|
||
slash command, a CLI one-liner), the script says it verbatim ("start your
|
||
prompt with /figma…") and the mock shows it being typed — the command is
|
||
the visible CAUSE of the result. Slash commands speak as "slash <name>",
|
||
caption as `/name`.
|
||
|
||
## Token-line format (`script-tokens.json`)
|
||
|
||
```json
|
||
{
|
||
"lines": [
|
||
{
|
||
"id": "l1",
|
||
"tokens": ["This", "week", "at", "HyperFrames,", { "display": "JSON", "spoken": "jay-sawn" }]
|
||
}
|
||
]
|
||
}
|
||
```
|
||
|
||
- A bare string = display and spoken identical.
|
||
- An object = the layers diverge. `display` keeps standard spelling AND the
|
||
punctuation captions should show; `spoken` is what the TTS reads.
|
||
- One line = one caption phrase (≤ ~40 chars of display text). Line grouping
|
||
is an authoring decision made here, not downstream.
|
||
- Build `vo-spoken.txt` by joining every token's spoken form with spaces,
|
||
lines joined into sentences/paragraphs as punctuated.
|
||
|
||
## Phonetics rules (ElevenLabs-style best practices, plain-text)
|
||
|
||
HeyGen TTS takes plain text (no SSML), so pronunciation is controlled by
|
||
spelling, hyphens, and spacing:
|
||
|
||
1. **Initialisms** (each letter said): space or hyphen the letters —
|
||
`CLI → "C L I"`, `CDP → "C D P"`, `API → "A P I"`.
|
||
2. **Acronyms said as words**: respell phonetically —
|
||
`JSON → "jay-sawn"`, `GSAP → "jee-sap"`.
|
||
3. **Mixed / pronounceable compounds**: hyphenated LOWERCASE phonetics, one
|
||
fluid run — `ffmpeg → "ff-mpeg"` (ear-tested; the TTS reads "ff" as a
|
||
fluid "eff-eff"), `WebM → "web em"`, `OAuth → "oh-auth"`. Never spaced
|
||
capitals here: the TTS reads spaced caps as isolated letter names with
|
||
hard stops ("F F em-peg" comes out "eff… eff… em-peg"). Reserve spaced
|
||
capitals for TRUE initialisms (CLI, API) where a deliberate
|
||
letter-by-letter read is the goal. When candidates are close, generate
|
||
A/B takes of the real sentence and let the user pick by ear.
|
||
4. **Versions/numbers**: expand — `v0.7.36 → "version zero point seven
|
||
point thirty-six"` (usually: don't speak versions at all),
|
||
`1080×1080 → "ten-eighty by ten-eighty"`.
|
||
5. **URLs**: `hyperframes.heygen.com → "hyperframes dot hey-jen dot com"`.
|
||
6. **Filenames/extensions**: `.mp4 → "dot em pee four"` — or rephrase so the
|
||
extension isn't spoken.
|
||
7. **Emphasis/pauses**: commas and em-dashes, never caps. Ellipses are
|
||
unreliable in TTS — use an em-dash.
|
||
|
||
The shared vocabulary lives in `references/lexicon.json`
|
||
(`display → spoken`). Consult it for EVERY technical term; if a term is
|
||
missing, ask the user for the pronunciation and add the entry — never guess,
|
||
never ship unheard. New entries: listen to that line in the generated VO
|
||
before accepting.
|
||
|
||
## Alignment (spoken timestamps → display captions)
|
||
|
||
`heygen-tts.mjs --words` returns word timestamps of the SPOKEN text.
|
||
`scripts/align-captions.mjs` walks the spoken stream against the token lines
|
||
(one display token may cover several spoken words — "C L I" is three) and
|
||
emits `captions.json`:
|
||
|
||
```json
|
||
{
|
||
"lines": [
|
||
{
|
||
"end": 3.1,
|
||
"w": [
|
||
["This", 0.22],
|
||
["week", 0.4],
|
||
["JSON", 1.1]
|
||
]
|
||
}
|
||
]
|
||
}
|
||
```
|
||
|
||
Each display word carries the start time of its FIRST spoken word; a line's
|
||
`end` = the next line's start (last line: last word end + 0.6). The aligner
|
||
warns `MISMATCH` when the heard word doesn't fuzzy-match the expected spoken
|
||
form — every warning must be resolved (fix the lexicon spelling or the
|
||
transcript) before the captions are trusted.
|
||
|
||
## Caption rail (rendering)
|
||
|
||
Per `captions-overlay`: a quiet OVERLAY, never a reserved band. One line,
|
||
bottom-center (top: 990px, height: 52px on 1080-square), TT Norms Pro 500 32px,
|
||
ink .94, soft dark text-shadow, words fading in (0.12s) on their timestamps,
|
||
phrase swaps as sets. Keep critical small text out of the bottom ~100px
|
||
center span; everything else may run under the rail.
|