Files
hyperframes/.claude/skills/changelog-video/references/script-voice.md
T
James RussoandJake Moran e96ebd74de feat(skills): add changelog-video skill for repo-native CC + Codex discovery (#2552)
Packages Jake Moran's changelog-video pipeline (v1, validated end-to-end
by Home on the Jun 23-29 range) as a repo-native skill set that Claude
Code (.claude/skills/) and Codex CLI (.agents/skills/) auto-discover the
moment the repo is opened. No install step; run the skill against a
changelog markdown for a given git range and it produces a lint-clean,
seam-gate-green 1080x1080 MP4 (~45-60s, Annie VO, mock-UI visualizations,
caption rail) end-to-end.

Six skills added byte-identical in both mirror dirs:
- changelog-video (pipeline entry point)
- motion-doctrine (carries seam-stamp.mjs + seam-gate.mjs)
- cut-the-curve, captions-overlay, seam-craft, oversized-cursor

Layout:
- .claude/skills/  - Claude Code project-local auto-discover
- .agents/skills/  - Codex CLI project-local auto-discover (verified via
                     Magi's clean-home Codex 0.144.3 repro; NOT .codex/skills/)

Fonts, animated background (12 MB), house BGM (5 MB), lexicon, and
align-captions ship inside the skill dirs. .gitattributes routes only
.claude/skills/**/*.{mp4,mp3} + .agents/skills/**/*.{mp4,mp3} through
LFS — narrowly scoped so unrelated Player, Studio, registry, and
marketplace media stay put. HeyGen CLI auth is the one credential the
skill needs; Node >= 22, ffmpeg, and headless Chrome are documented
alongside in both READMEs.

.gitignore: rewrites .claude/ and .agents/ blocks to keep agent-installed
skill hygiene while re-including the six repo-native skill dirs plus
README.md.

CI:
- Extends changes.skills filter to match .claude/skills/**,
  .agents/skills/**, scripts/lint-skills.ts, and scripts/check-skill-mirror.mjs.
- New 'Skills: project-native lint + mirror' job runs the extended
  lint-skills.ts (schema-driven; required { name, description } + optional
  { license, allowed-tools, metadata }, name pattern check, description
  length check) plus a new check-skill-mirror.mjs byte-integrity script
  (24 mirrored files must match; README.md deliberately per-CLI).
- Wired into 'bun run lint' locally.

Frontmatter validator:
- Rejects unsupported top-level keys (catches category:-style drift).
- Requires name + description.
- Validates name pattern (^[a-z][a-z0-9-]{0,63}$) and description shape
  (non-empty, <=1024 chars).
- Missing frontmatter block itself is a first-class error.

Also strips unsupported top-level 'category:' frontmatter from Jake's
motion-doctrine and cut-the-curve SKILL.mds (both mirrors), rewrites the
TTS invocation from ~/.claude/skills/media-use/... to the tracked
skills/hyperframes-media/scripts/heygen-tts.mjs, swaps npx hyperframes@latest
for the repo-local CLI in the gate step, and fixes a lint issue in Jake's
seam-gate.mjs (ternary-for-side-effect -> if/else).

Validated end-to-end by Home on Jun 23-29 (MP4 posted in C0ACCNHLG3U
thread 1784181166.041319). Independently reviewed R1/R2/R3 by Magi.

Co-authored-by: Jake Moran <jake@heygen.com>
2026-07-16 17:29:19 -04:00

112 lines
4.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Script + Voice: the two-layer contract
The script is the single source of truth for BOTH the VO and the captions,
written as token lines. The VO reads the `spoken` layer; captions render the
`display` layer. This is a hard quality gate: a caption showing "jay-sawn" or
a VO saying "juh-son" ("JSON" read literally) are both build failures.
## Register (how it should sound)
- Conversational, not release-notes. "The big one this week —" beats
"Theme 1:". Contractions welcome. Second person allowed ("your clips").
- Informational, never salesy; no superlatives the changelog doesn't earn.
- One breath per beat: sentences ≤ ~14 words; let punctuation pace the read.
- Numbers with meaning stay ("fifteen releases"); commit hashes, PR numbers,
and version micro-detail are never spoken.
- Open with the week + the marquee, close with the digest pointer
("See everything at hyperframes dot heygen dot com").
- **Teach the simple command.** When a feature has a one-line invocation (a
slash command, a CLI one-liner), the script says it verbatim ("start your
prompt with /figma…") and the mock shows it being typed — the command is
the visible CAUSE of the result. Slash commands speak as "slash <name>",
caption as `/name`.
## Token-line format (`script-tokens.json`)
```json
{
"lines": [
{
"id": "l1",
"tokens": ["This", "week", "at", "HyperFrames,", { "display": "JSON", "spoken": "jay-sawn" }]
}
]
}
```
- A bare string = display and spoken identical.
- An object = the layers diverge. `display` keeps standard spelling AND the
punctuation captions should show; `spoken` is what the TTS reads.
- One line = one caption phrase (≤ ~40 chars of display text). Line grouping
is an authoring decision made here, not downstream.
- Build `vo-spoken.txt` by joining every token's spoken form with spaces,
lines joined into sentences/paragraphs as punctuated.
## Phonetics rules (ElevenLabs-style best practices, plain-text)
HeyGen TTS takes plain text (no SSML), so pronunciation is controlled by
spelling, hyphens, and spacing:
1. **Initialisms** (each letter said): space or hyphen the letters —
`CLI → "C L I"`, `CDP → "C D P"`, `API → "A P I"`.
2. **Acronyms said as words**: respell phonetically —
`JSON → "jay-sawn"`, `GSAP → "jee-sap"`.
3. **Mixed / pronounceable compounds**: hyphenated LOWERCASE phonetics, one
fluid run — `ffmpeg → "ff-mpeg"` (ear-tested; the TTS reads "ff" as a
fluid "eff-eff"), `WebM → "web em"`, `OAuth → "oh-auth"`. Never spaced
capitals here: the TTS reads spaced caps as isolated letter names with
hard stops ("F F em-peg" comes out "eff… eff… em-peg"). Reserve spaced
capitals for TRUE initialisms (CLI, API) where a deliberate
letter-by-letter read is the goal. When candidates are close, generate
A/B takes of the real sentence and let the user pick by ear.
4. **Versions/numbers**: expand — `v0.7.36 → "version zero point seven
point thirty-six"` (usually: don't speak versions at all),
`1080×1080 → "ten-eighty by ten-eighty"`.
5. **URLs**: `hyperframes.heygen.com → "hyperframes dot hey-jen dot com"`.
6. **Filenames/extensions**: `.mp4 → "dot em pee four"` — or rephrase so the
extension isn't spoken.
7. **Emphasis/pauses**: commas and em-dashes, never caps. Ellipses are
unreliable in TTS — use an em-dash.
The shared vocabulary lives in `references/lexicon.json`
(`display → spoken`). Consult it for EVERY technical term; if a term is
missing, ask the user for the pronunciation and add the entry — never guess,
never ship unheard. New entries: listen to that line in the generated VO
before accepting.
## Alignment (spoken timestamps → display captions)
`heygen-tts.mjs --words` returns word timestamps of the SPOKEN text.
`scripts/align-captions.mjs` walks the spoken stream against the token lines
(one display token may cover several spoken words — "C L I" is three) and
emits `captions.json`:
```json
{
"lines": [
{
"end": 3.1,
"w": [
["This", 0.22],
["week", 0.4],
["JSON", 1.1]
]
}
]
}
```
Each display word carries the start time of its FIRST spoken word; a line's
`end` = the next line's start (last line: last word end + 0.6). The aligner
warns `MISMATCH` when the heard word doesn't fuzzy-match the expected spoken
form — every warning must be resolved (fix the lexicon spelling or the
transcript) before the captions are trusted.
## Caption rail (rendering)
Per `captions-overlay`: a quiet OVERLAY, never a reserved band. One line,
bottom-center (top: 1002px on 1080-square), TT Norms Pro 500 25px,
ink .92, soft dark text-shadow, words fading in (0.12s) on their timestamps,
phrase swaps as sets. Keep critical small text out of the bottom ~80px
center span; everything else may run under the rail.