mirror of
https://github.com/heygen-com/hyperframes.git
synced 2026-09-13 15:49:53 +00:00
fix(skills,cli): close four reproduced contract gaps from the CLI feedback digest (#2476)
* fix(cli): invalidate the skills nudge cache after a successful install/update/check The passive "N skills out of date or missing" nudge reads a 24h config cache that only the background check (on non-skills commands) ever wrote. The skills commands themselves are excluded from the nudge pipeline, so a successful `skills update`/install/check never refreshed or dropped the cached verdict — the pre-install count kept printing on every other command for up to 24h. Reconcile commands now drop the cached verdict (counts + timestamp) so the next command's background check re-runs for real. The offline presence-only path deliberately keeps the cache: that run learned nothing about freshness. * fix(skills): win32-safe npx spawns in media-use + accurate whisper wording The Whisper transcribe fallback and the Kokoro local-TTS delegation both spawned a bare "npx" via execFileSync — on Windows npx is npx.cmd, which spawn cannot exec, so both paths died with `spawnSync npx ENOENT`. Route them through the skill's existing resolveSpawnCommand (node + npx-cli.js on win32, no shell:true), same as the audio engine's TTS spawns. Also corrects the "bundled with the hyperframes CLI" claim about whisper.cpp: it is resolved from PATH / installed via Homebrew / built from source with git+cmake on first use, and models download from HuggingFace — nothing whisper is shipped in the package. * feat(skills): canonical fully-silent marker + auth status exit-code docs product-launch's Step 3.1 gate said "or the project is marked silent" but nothing defined how to mark one, and audio.mjs unconditionally retrieved BGM. Define the canonical marker — `music: none` in the storyboard's top YAML block, plus no SCRIPT.md — and honor it: audio generate produces nothing (removing stale audio_meta.json, since absence is what assemble treats as silent), and `music: none` with narration keeps TTS while turning BGM off. Also documents the `auth status` exit-code contract (exit 1 while signed out is the normal offline state, not a failure) in the product-launch Step 0 note and the CLI skill's cloud reference. * fix(skills): transient-init retry for standalone animation-map and contrast-report The standalone helpers called initializeSession exactly once, so a valid modular project — whose sub-composition timelines register asynchronously — could hit the readiness deadline and die with the transient "zero duration / Runtime ready: false" diagnostic the render pipeline retries (probeStage). Add initializeSessionWithRetry to the shared package-loader (both byte-identical copies): close the crashed session and retry once with a fresh browser, gated by the engine's canonical isTransientBrowserError — now re-exported from @hyperframes/producer, with a frozen fallback pattern list for older published packages. The "Runtime ready: true" fast-fail (a genuine authoring bug) still fails without a retry. * feat(skills): extend the fully-silent marker to faceless-explainer and pr-to-video Both workflows reuse product-launch's audio model — their Step 3.1 gates carried the same undefined "marked silent" phrase, and their (intentionally identical) audio.mjs copies had the same unconditional BGM retrieve. Port the `music: none` marker handling into both copies, define the marker in their SKILL.md Step 3.1 and story-design references, and turn the copies' "intentionally identical" header claim into a byte-identity pin test so the next fix can't silently miss one of them. * test(cli): reset the prune mock explicitly instead of relying on restoreAllMocks The converge test's toHaveBeenCalledTimes(1) held only because vitest 3's vi.restoreAllMocks() clears vi.fn() call state; vitest 4 restores spies only, so the count would accumulate across tests and fail. Reset pruneOrphanedLockEntries in beforeEach like the other manifest mocks — passes under both vitest 3.2.4 (pinned) and vitest 4. * test(skills): close review findings — package-loader pin, whisper win32 parity, quoted-none Review follow-ups on #2476: - package-loader.mjs byte-identity pin (the elevated concern): the two copies now carry initializeSessionWithRetry + FALLBACK_TRANSIENT_PATTERNS, exactly the shared-logic shape a future fix could land in one copy and miss in the other — same enforcement as the audio.mjs pin. - whisper win32 call-site parity: runWhisper's npx resolution lifted into lib/npx-sync.mjs (resolveNpxInvocation, injectable params matching the localTtsGenerate idiom) with the same three-branch coverage as the Kokoro site — plus the hard-fail contract (throws actionably, since the whisper fallback has no next provider to fall through to). - quoted music: "none" pin: the vendored storyboard parser strips matching quotes at parse time (stripQuotes), so the silent marker already accepts the quoted spelling — pinned so that stays true.
This commit is contained in:
@@ -35,7 +35,7 @@ After init, let `<PROJECT_ROOT>` be `videos/<project>` and run every subsequent
|
||||
|
||||
**Write `BRIEF.md` immediately after init** (never before — `init` refuses a non-empty directory): the intent layer's locked brief, shape per `../hyperframes-core/references/brief-format.md`. Resolve `<MEDIA_DIR>` as the installed `/media-use` skill directory. Then record each preference-backed answer with `node <MEDIA_DIR>/scripts/prefs.mjs record --hyperframes .` (`brief-format.md` names the subset). If the intent layer adopted a recipe, run `node <MEDIA_DIR>/scripts/recipe.mjs use --hyperframes . --name <name>`; it copies its `frame.md` into the project (Step 2 is then skipped) and returns the skeletons Step 3 drafts from. A recipe fills answers, not approvals; the review gates still run.
|
||||
|
||||
**Show sign-in status before proceeding past Setup** — run `npx hyperframes auth status` and relay its output verbatim. It reports whether voice/BGM will use HeyGen or local engines and, when signed out, how to sign in. Apply one branch:
|
||||
**Show sign-in status before proceeding past Setup** — run `npx hyperframes auth status` and relay its output verbatim. It reports whether voice/BGM will use HeyGen or local engines and, when signed out, how to sign in. Note the exit code contract: `auth status` **exits 1 when not signed in** (and when the stored credential is rejected) — that non-zero exit is the normal signed-out state, not a command failure, so don't treat it as an error, don't retry it, and don't chain it with `&&`/`set -e` in a way that would abort the workflow. Apply one branch:
|
||||
|
||||
- **Collaborative:** wait for the user to sign in or explicitly choose `offline` / `go`.
|
||||
- **Autonomous:** state the status and continue through the available local engines.
|
||||
@@ -104,11 +104,13 @@ Start audio after Step 3 approval. Run it in the background, then continue to St
|
||||
|
||||
`node <SKILL_DIR>/scripts/audio.mjs --script ./SCRIPT.md --storyboard ./STORYBOARD.md --hyperframes . --out ./audio_meta.json --provider <provider> --voice <voice-id> &`
|
||||
|
||||
The audio script handles narration, word timings, BGM lookup from HeyGen's music library, and timing metadata. BGM mood comes from the storyboard's `music:` field. This uses the HeyGen Audio API for retrieval, not generation, and uses the same `~/.heygen` credential as TTS. For provider details, read `../media-use/audio/references/tts.md`.
|
||||
The audio script handles narration, word timings, BGM lookup from HeyGen's music library, and timing metadata. BGM mood comes from the storyboard's `music:` field; **`music: none` turns BGM off**. This uses the HeyGen Audio API for retrieval, not generation, and uses the same `~/.heygen` credential as TTS. For provider details, read `../media-use/audio/references/tts.md`.
|
||||
|
||||
If there is no narration and no `SCRIPT.md`, skip voice generation. BGM may still run if the storyboard has a music mood.
|
||||
|
||||
**Gate:** audio job has started, or the project is marked silent.
|
||||
**The canonical fully-silent marker:** `music: none` in the STORYBOARD.md top YAML block **and** no `SCRIPT.md`. That combination marks the project silent — no narration, no BGM, no SFX. `audio.mjs` recognizes it and generates nothing (it removes any stale `audio_meta.json`; an absent `audio_meta.json` is what assemble treats as silent), so Step 3.1 is a clean skip. Use it when the user asks for a silent / music-free video — don't improvise other spellings.
|
||||
|
||||
**Gate:** audio job has started, or the project is marked silent (`music: none` + no `SCRIPT.md`).
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -332,6 +332,13 @@ The five registry types:
|
||||
|
||||
Pick a small set and repeat them: default to `crossfade` (or `blur-crossfade` when the backgrounds clash), and reach for `zoom-through` at section boundaries. Frame 1's `transition_in` is a placeholder.
|
||||
|
||||
## Music & silence
|
||||
|
||||
The storyboard's top YAML block carries a `music:` field — the BGM mood the audio step retrieves against (e.g. `music: confident minimal tech underscore`). Omitting it falls back to `message:` → `arc:` → a neutral default, so BGM plays unless turned off explicitly.
|
||||
|
||||
- **`music: none`** — BGM off (narration, if any, still runs).
|
||||
- **`music: none` + no `SCRIPT.md`** — the canonical **fully-silent marker**: no narration, no BGM, no SFX. `audio.mjs` generates nothing and Step 3.1 is a clean skip. Use exactly this spelling when the user asks for a silent / music-free video.
|
||||
|
||||
## Frame template
|
||||
|
||||
Use the exact fields required by the core storyboard format. The narrative shape each frame satisfies:
|
||||
|
||||
@@ -21,7 +21,7 @@
|
||||
// node audio.mjs fetch-sfx --storyboard ./STORYBOARD.md --hyperframes .
|
||||
|
||||
import { spawnSync } from "node:child_process";
|
||||
import { existsSync, readFileSync, writeFileSync } from "node:fs";
|
||||
import { existsSync, readFileSync, rmSync, writeFileSync } from "node:fs";
|
||||
import { dirname, join, resolve } from "node:path";
|
||||
import { fileURLToPath } from "node:url";
|
||||
import { parseStoryboard } from "./lib/storyboard.mjs";
|
||||
@@ -137,6 +137,22 @@ function runGenerate(argv) {
|
||||
text: l.text,
|
||||
}))
|
||||
: [];
|
||||
// The canonical fully-silent marker (SKILL.md Step 3.1): `music: none` in
|
||||
// the storyboard's top YAML block turns BGM off; combined with no SCRIPT.md
|
||||
// the project is fully silent — generate nothing and remove any stale meta
|
||||
// from a previous run (assemble treats an absent audio_meta.json as silent).
|
||||
const bgmOff =
|
||||
String(g.extra?.music ?? "")
|
||||
.trim()
|
||||
.toLowerCase() === "none";
|
||||
if (bgmOff && !lines.length) {
|
||||
rmSync(outPath, { force: true });
|
||||
rmSync(neutralPath(outPath), { force: true });
|
||||
console.log(
|
||||
"✓ audio generate: project marked silent (music: none, no SCRIPT.md) — nothing to generate",
|
||||
);
|
||||
return;
|
||||
}
|
||||
if (!lines.length) console.error("· no SCRIPT.md — silent film (BGM only)");
|
||||
|
||||
// BGM mood: storyboard `music:` → message → arc → default. `mode: retrieve` is
|
||||
@@ -146,7 +162,9 @@ function runGenerate(argv) {
|
||||
provider,
|
||||
speed,
|
||||
lines,
|
||||
bgm: { mode: "retrieve", query, blob: g.message || "", arc: g.arc || "" },
|
||||
bgm: bgmOff
|
||||
? { mode: "none" }
|
||||
: { mode: "retrieve", query, blob: g.message || "", arc: g.arc || "" },
|
||||
};
|
||||
if (userVoice) request.voice = userVoice;
|
||||
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
import assert from "node:assert/strict";
|
||||
import { mkdtempSync, readFileSync, writeFileSync } from "node:fs";
|
||||
import { existsSync, mkdtempSync, readFileSync, writeFileSync } from "node:fs";
|
||||
import { tmpdir } from "node:os";
|
||||
import { join } from "node:path";
|
||||
import { spawnSync } from "node:child_process";
|
||||
@@ -44,3 +44,90 @@ test("--provider takes precedence over HF_TTS_PROVIDER", () => {
|
||||
"kokoro",
|
||||
);
|
||||
});
|
||||
|
||||
// ── the canonical fully-silent marker (SKILL.md Step 3.1) ────────────────────
|
||||
// `music: none` in the storyboard's top YAML block + no SCRIPT.md marks the
|
||||
// project fully silent: generate must produce nothing (an absent
|
||||
// audio_meta.json is what assemble treats as silent) and clear stale meta.
|
||||
|
||||
/** Like runAudio, but with a caller-controlled storyboard and no request assertion. */
|
||||
function runAudioRaw({ storyboard, scriptMd = null, preexistingMeta = null }) {
|
||||
const dir = mkdtempSync(join(tmpdir(), "product-launch-audio-"));
|
||||
const engine = join(dir, "engine.mjs");
|
||||
writeFileSync(join(dir, "STORYBOARD.md"), storyboard);
|
||||
if (scriptMd != null) writeFileSync(join(dir, "SCRIPT.md"), scriptMd);
|
||||
if (preexistingMeta != null) writeFileSync(join(dir, "audio_meta.json"), preexistingMeta);
|
||||
writeFileSync(
|
||||
engine,
|
||||
`import { readFileSync, writeFileSync } from "node:fs";
|
||||
const argv = process.argv.slice(2);
|
||||
const flag = (name) => argv[argv.indexOf(name) + 1];
|
||||
const request = JSON.parse(readFileSync(flag("--request"), "utf8"));
|
||||
writeFileSync(new URL("request.json", import.meta.url), JSON.stringify(request));
|
||||
writeFileSync(flag("--out"), JSON.stringify({ voices: [], bgm: null, sfx: [] }));
|
||||
`,
|
||||
);
|
||||
const result = spawnSync(
|
||||
process.execPath,
|
||||
[script, "--hyperframes", dir, "--storyboard", join(dir, "STORYBOARD.md")],
|
||||
{ encoding: "utf8", env: { ...process.env, HF_MEDIA_ENGINE: engine } },
|
||||
);
|
||||
return { dir, result };
|
||||
}
|
||||
|
||||
test("music: none + no SCRIPT.md = fully silent: no engine run, no audio_meta.json", () => {
|
||||
const { dir, result } = runAudioRaw({ storyboard: "---\nmessage: Test\nmusic: none\n---\n" });
|
||||
|
||||
assert.equal(result.status, 0, result.stderr);
|
||||
assert.match(result.stdout, /marked silent/);
|
||||
// The engine was never invoked...
|
||||
assert.equal(existsSync(join(dir, "request.json")), false);
|
||||
// ...and no meta exists (absence ⇒ assemble treats the film as silent).
|
||||
assert.equal(existsSync(join(dir, "audio_meta.json")), false);
|
||||
});
|
||||
|
||||
test("fully-silent run removes stale audio_meta.json from a previous non-silent run", () => {
|
||||
const { dir, result } = runAudioRaw({
|
||||
storyboard: "---\nmessage: Test\nmusic: none\n---\n",
|
||||
preexistingMeta: JSON.stringify({ bgm: { path: "old.mp3" }, voices: [], sfx: [] }),
|
||||
});
|
||||
|
||||
assert.equal(result.status, 0, result.stderr);
|
||||
assert.equal(existsSync(join(dir, "audio_meta.json")), false);
|
||||
});
|
||||
|
||||
test("music: none with narration keeps TTS but turns BGM off (not fully silent)", () => {
|
||||
const { dir, result } = runAudioRaw({
|
||||
storyboard: "---\nmessage: Test\nmusic: none\n---\n",
|
||||
scriptMd: "## Hook (Frame 1)\n\n Spoken line for frame one.\n",
|
||||
});
|
||||
|
||||
assert.equal(result.status, 0, result.stderr);
|
||||
const request = JSON.parse(readFileSync(join(dir, "request.json"), "utf8"));
|
||||
assert.equal(request.bgm.mode, "none");
|
||||
assert.equal(request.lines.length, 1);
|
||||
});
|
||||
|
||||
test('a quoted music: "none" is still the silent marker (frontmatter stripQuotes)', () => {
|
||||
// YAML authors quote scalars freely; the vendored storyboard parser strips
|
||||
// matching quotes at parse time (storyboard.mjs stripQuotes), so the marker
|
||||
// must not depend on the unquoted spelling.
|
||||
const { dir, result } = runAudioRaw({
|
||||
storyboard: '---\nmessage: Test\nmusic: "none"\n---\n',
|
||||
});
|
||||
|
||||
assert.equal(result.status, 0, result.stderr);
|
||||
assert.match(result.stdout, /marked silent/);
|
||||
assert.equal(existsSync(join(dir, "audio_meta.json")), false);
|
||||
});
|
||||
|
||||
test("a storyboard music mood still retrieves BGM (marker is exact, not fuzzy)", () => {
|
||||
const { dir, result } = runAudioRaw({
|
||||
storyboard: "---\nmessage: Test\nmusic: upbeat synthwave with heavy drums\n---\n",
|
||||
});
|
||||
|
||||
assert.equal(result.status, 0, result.stderr);
|
||||
const request = JSON.parse(readFileSync(join(dir, "request.json"), "utf8"));
|
||||
assert.equal(request.bgm.mode, "retrieve");
|
||||
assert.equal(request.bgm.query, "upbeat synthwave with heavy drums");
|
||||
});
|
||||
|
||||
Reference in New Issue
Block a user