* fix(cli): invalidate the skills nudge cache after a successful install/update/check The passive "N skills out of date or missing" nudge reads a 24h config cache that only the background check (on non-skills commands) ever wrote. The skills commands themselves are excluded from the nudge pipeline, so a successful `skills update`/install/check never refreshed or dropped the cached verdict — the pre-install count kept printing on every other command for up to 24h. Reconcile commands now drop the cached verdict (counts + timestamp) so the next command's background check re-runs for real. The offline presence-only path deliberately keeps the cache: that run learned nothing about freshness. * fix(skills): win32-safe npx spawns in media-use + accurate whisper wording The Whisper transcribe fallback and the Kokoro local-TTS delegation both spawned a bare "npx" via execFileSync — on Windows npx is npx.cmd, which spawn cannot exec, so both paths died with `spawnSync npx ENOENT`. Route them through the skill's existing resolveSpawnCommand (node + npx-cli.js on win32, no shell:true), same as the audio engine's TTS spawns. Also corrects the "bundled with the hyperframes CLI" claim about whisper.cpp: it is resolved from PATH / installed via Homebrew / built from source with git+cmake on first use, and models download from HuggingFace — nothing whisper is shipped in the package. * feat(skills): canonical fully-silent marker + auth status exit-code docs product-launch's Step 3.1 gate said "or the project is marked silent" but nothing defined how to mark one, and audio.mjs unconditionally retrieved BGM. Define the canonical marker — `music: none` in the storyboard's top YAML block, plus no SCRIPT.md — and honor it: audio generate produces nothing (removing stale audio_meta.json, since absence is what assemble treats as silent), and `music: none` with narration keeps TTS while turning BGM off. Also documents the `auth status` exit-code contract (exit 1 while signed out is the normal offline state, not a failure) in the product-launch Step 0 note and the CLI skill's cloud reference. * fix(skills): transient-init retry for standalone animation-map and contrast-report The standalone helpers called initializeSession exactly once, so a valid modular project — whose sub-composition timelines register asynchronously — could hit the readiness deadline and die with the transient "zero duration / Runtime ready: false" diagnostic the render pipeline retries (probeStage). Add initializeSessionWithRetry to the shared package-loader (both byte-identical copies): close the crashed session and retry once with a fresh browser, gated by the engine's canonical isTransientBrowserError — now re-exported from @hyperframes/producer, with a frozen fallback pattern list for older published packages. The "Runtime ready: true" fast-fail (a genuine authoring bug) still fails without a retry. * feat(skills): extend the fully-silent marker to faceless-explainer and pr-to-video Both workflows reuse product-launch's audio model — their Step 3.1 gates carried the same undefined "marked silent" phrase, and their (intentionally identical) audio.mjs copies had the same unconditional BGM retrieve. Port the `music: none` marker handling into both copies, define the marker in their SKILL.md Step 3.1 and story-design references, and turn the copies' "intentionally identical" header claim into a byte-identity pin test so the next fix can't silently miss one of them. * test(cli): reset the prune mock explicitly instead of relying on restoreAllMocks The converge test's toHaveBeenCalledTimes(1) held only because vitest 3's vi.restoreAllMocks() clears vi.fn() call state; vitest 4 restores spies only, so the count would accumulate across tests and fail. Reset pruneOrphanedLockEntries in beforeEach like the other manifest mocks — passes under both vitest 3.2.4 (pinned) and vitest 4. * test(skills): close review findings — package-loader pin, whisper win32 parity, quoted-none Review follow-ups on #2476: - package-loader.mjs byte-identity pin (the elevated concern): the two copies now carry initializeSessionWithRetry + FALLBACK_TRANSIENT_PATTERNS, exactly the shared-logic shape a future fix could land in one copy and miss in the other — same enforcement as the audio.mjs pin. - whisper win32 call-site parity: runWhisper's npx resolution lifted into lib/npx-sync.mjs (resolveNpxInvocation, injectable params matching the localTtsGenerate idiom) with the same three-branch coverage as the Kokoro site — plus the hard-fail contract (throws actionably, since the whisper fallback has no next provider to fall through to). - quoted music: "none" pin: the vendored storyboard parser strips matching quotes at parse time (stripQuotes), so the silent marker already accepts the quoted spelling — pinned so that stays true.
4.2 KiB
Requirements & Caches
Credential & key priority
Run npx hyperframes auth status to see what's configured and which engines a workflow will use (see the skill's Preflight section). Keys resolve in this order — first match wins:
| Provider | Resolution order (first non-empty wins) | Local deps when used |
|---|---|---|
| HeyGen (TTS + BGM/SFX retrieval) | $HEYGEN_API_KEY → $HYPERFRAMES_API_KEY → ~/.heygen/credentials (shared with heygen-cli; $HEYGEN_CONFIG_DIR overrides the dir; written by hyperframes auth login) |
none (REST) |
| ElevenLabs (TTS fallback) | $ELEVENLABS_API_KEY |
pip install elevenlabs |
| Lyria (BGM fallback) | $GEMINI_API_KEY → $GOOGLE_API_KEY |
pip install google-genai |
| Kokoro (TTS, no key) | always — final voice fallback | pip install kokoro-onnx soundfile |
| MusicGen (BGM, no key) | always — final music fallback | pip install transformers torch soundfile numpy |
hyperframes auth login (browser OAuth) is the recommended setup: one sign-in, every project, no per-repo .env. An OAuth login is sent as Authorization: Bearer; an API key as X-Api-Key; both are tagged with X-HeyGen-Source: cli. OAuth CLI users can consume the web-plan free allowance for HeyGen TTS (10 min/month); API keys follow the normal API billing path. With no HeyGen credential, voice/BGM run fully locally (Kokoro / MusicGen) — hyperframes auth status and hyperframes doctor both report whether those local deps are installed.
Model caches & system dependencies
Each command downloads its own model on first run and caches it under ~/.cache/hyperframes/:
- TTS (HeyGen) — no local deps; needs a HeyGen credential +
ffmpegon PATH (to transcode the mp3 response to.wav). Credential resolves like the CLI:$HEYGEN_API_KEY→$HYPERFRAMES_API_KEY→~/.heygen/credentials(shared with heygen-cli; runnpx hyperframes auth login). An OAuth login is sent asAuthorization: Bearer; an API key asX-Api-Key; both includeX-HeyGen-Source: cliso the backend can apply CLI OAuth free usage. - TTS (ElevenLabs) — same as HeyGen: API key +
ffmpeg. - TTS (Kokoro) — Kokoro-82M (~311 MB) + voices (~27 MB) in
tts/. Requires Python 3.8+ withkokoro-onnxandsoundfile(pip install kokoro-onnx soundfile). Non-English text also needsespeak-ngsystem-wide. - BGM (Lyria) — needs
$GEMINI_API_KEYor$GOOGLE_API_KEY+pip install google-genai. No local model cache. - BGM (MusicGen) —
pip install transformers torch soundfile.facebook/musicgen-small(~300 MB) cached under~/.cache/huggingface/on first run. - Transcribe — Whisper model size depending on choice (75 MB – 3.1 GB) in
whisper/, downloaded from HuggingFace on first use.whisper.cppitself is NOT bundled: the CLI resolves it from PATH, installs via Homebrew (macOS), or builds it from source with git+cmake on first use ($HYPERFRAMES_WHISPER_PATHoverrides). - Remove-background —
u2net_human_seg(~168 MB ONNX) inbackground-removal/models/. Peak inference RAM ~1.5 GB.
Run npx hyperframes doctor if a command fails because of a missing dependency.