Code review found the new hyperframes-media skill was parallel content with skills/hyperframes/references/tts.md and the "Whisper Model Guide" section of transcript-guide.md — same voice table, same .en-translates-non-English warning, same TTS→transcribe chain in both places. Plus some scope creep in hyperframes-media (audio/video HTML snippets that duplicate the canonical track docs in hyperframes/SKILL.md:265+). Consolidation: - hyperframes-media is now the single source of truth for CLI invocation, voice selection, multilingual phonemization, whisper model selection, and the .en gotcha. Picked up the multilingual prefix decoding from the deleted tts.md. - skills/hyperframes/references/tts.md deleted; the bullet in hyperframes/SKILL.md is removed (no replacement — agents land on hyperframes-media via its own description). - skills/hyperframes/references/transcript-guide.md keeps only the caption-side concerns: input-format table, mandatory quality check, cleaning JS, external-API import path, and the "if no transcript exists" flow. The intro bash recipe and Whisper Model Guide section both moved to hyperframes-media. Top of the file now points to hyperframes-media for CLI/model details. Other tightening in hyperframes-media: - Dropped WHAT-narration filler and the inline <audio>/<video> HTML snippets — they duplicate the canonical track-attribute docs in hyperframes/SKILL.md. - Added the `id` field (`w0`, `w1`, ...) to the transcript output shape — the actual Word interface in packages/cli/src/whisper/normalize.ts includes it (optional for backwards compat), used by caption override logic. - Compressed the TTS → Transcribe → Captions chain section. Net: hyperframes-media 147 → 136 lines, transcript-guide.md 152 → 106 lines, tts.md gone (-75 lines).
5.7 KiB
Transcript Guide
For the transcribe CLI invocation, the .en-translates-non-English rule, and whisper model selection, see the hyperframes-media skill. This file covers what to do with the resulting transcript when authoring captions: input formats, mandatory quality checks, cleaning code, external-API fallbacks.
Supported Input Formats
The CLI auto-detects and normalizes these formats:
| Format | Extension | Source | Word-level? |
|---|---|---|---|
| whisper.cpp JSON | .json |
hyperframes init --video, hyperframes transcribe |
Yes |
| OpenAI Whisper API | .json |
openai.audio.transcriptions.create({ timestamp_granularities: ["word"] }) |
Yes |
| SRT subtitles | .srt |
Video editors, subtitle tools, YouTube | No (phrase-level) |
| VTT subtitles | .vtt |
Web players, YouTube, transcription services | No (phrase-level) |
| Normalized word array | .json |
Pre-processed by any tool | Yes |
Word-level timestamps produce better captions. SRT/VTT give phrase-level timing, which works but can't do per-word animation effects.
Transcript Quality Check (Mandatory)
After every transcription, read the transcript and check for quality issues before proceeding. Bad transcripts produce nonsensical captions. Never skip this step.
What to look for
| Signal | Example | Cause |
|---|---|---|
Music note tokens (♪, �) |
{ "text": "♪" } or { "text": "�" } |
Whisper detected music, not speech |
| Garbled / nonsense words | "Do a chin", "Get so gay", "huh" | Model misheard lyrics or background noise |
| Long gaps with no words | 20+ seconds of only ♪ tokens |
Instrumental section — expected, but high ratio means speech is being missed |
| Repeated filler | Many "huh", "uh", "oh" entries | Model is hallucinating on music |
| Very short word spans | Words with end - start < 0.05 |
Unreliable timestamp alignment |
Automatic retry rules
If more than 20% of entries are ♪/� tokens, or the transcript contains obvious nonsense words, the transcription failed. Do not proceed with the bad transcript. Instead:
- Retry with
medium.enif the original usedsmall.enor smaller:npx hyperframes transcribe audio.mp3 --model medium.en - If
medium.enalso fails (still >20% music tokens or garbled), tell the user the audio is too noisy for local transcription and suggest:- Providing lyrics manually as an SRT/VTT file
- Using an external API (OpenAI or Groq Whisper — see below)
- Always clean the transcript before building captions — filter out
♪/�tokens and entries wheretextis a single non-word character. Only real words should reach the caption composition.
Cleaning a transcript
After transcription (even with a good model), strip non-word entries:
var raw = JSON.parse(transcriptJson);
var words = raw.filter(function (w) {
if (!w.text || w.text.trim().length === 0) return false;
if (/^[♪�\u266a\u266b\u266c\u266d\u266e\u266f]+$/.test(w.text)) return false;
if (/^(huh|uh|um|ah|oh)$/i.test(w.text) && w.end - w.start < 0.1) return false;
return true;
});
When to use which model (decision tree)
- Is this speech over silence/light background? →
small.enis fine - Is this speech over music, or music with vocals? → Start with
medium.en - Is this a produced music track (vocals + full instrumentation)? → Start with
medium.en, expect to need manual lyrics or an external API - Is this multilingual? → Use
mediumorlarge-v3(no.ensuffix)
Using External Transcription APIs
For the best accuracy, use an external API and import the result:
OpenAI Whisper API (recommended for quality):
# Generate with word timestamps, then import
curl https://api.openai.com/v1/audio/transcriptions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-F file=@audio.mp3 -F model=whisper-1 \
-F response_format=verbose_json \
-F "timestamp_granularities[]=word" \
-o transcript-openai.json
npx hyperframes transcribe transcript-openai.json
Groq Whisper API (fast, free tier available):
curl https://api.groq.com/openai/v1/audio/transcriptions \
-H "Authorization: Bearer $GROQ_API_KEY" \
-F file=@audio.mp3 -F model=whisper-large-v3 \
-F response_format=verbose_json \
-F "timestamp_granularities[]=word" \
-o transcript-groq.json
npx hyperframes transcribe transcript-groq.json
If No Transcript Exists
- Check the project root for
transcript.json,.srt, or.vttfiles - If none found, run transcription — pick the starting model based on the content type:
- Speech/voiceover →
small.en - Music with vocals →
medium.en
npx hyperframes transcribe <audio-or-video-file> --model medium.en - Speech/voiceover →
- Read the transcript and run the quality check (see above). If it fails, retry with a larger model or suggest manual lyrics.