mirror of
https://github.com/heygen-com/hyperframes.git
synced 2026-09-11 06:30:03 +00:00
* fix(cli): invalidate the skills nudge cache after a successful install/update/check The passive "N skills out of date or missing" nudge reads a 24h config cache that only the background check (on non-skills commands) ever wrote. The skills commands themselves are excluded from the nudge pipeline, so a successful `skills update`/install/check never refreshed or dropped the cached verdict — the pre-install count kept printing on every other command for up to 24h. Reconcile commands now drop the cached verdict (counts + timestamp) so the next command's background check re-runs for real. The offline presence-only path deliberately keeps the cache: that run learned nothing about freshness. * fix(skills): win32-safe npx spawns in media-use + accurate whisper wording The Whisper transcribe fallback and the Kokoro local-TTS delegation both spawned a bare "npx" via execFileSync — on Windows npx is npx.cmd, which spawn cannot exec, so both paths died with `spawnSync npx ENOENT`. Route them through the skill's existing resolveSpawnCommand (node + npx-cli.js on win32, no shell:true), same as the audio engine's TTS spawns. Also corrects the "bundled with the hyperframes CLI" claim about whisper.cpp: it is resolved from PATH / installed via Homebrew / built from source with git+cmake on first use, and models download from HuggingFace — nothing whisper is shipped in the package. * feat(skills): canonical fully-silent marker + auth status exit-code docs product-launch's Step 3.1 gate said "or the project is marked silent" but nothing defined how to mark one, and audio.mjs unconditionally retrieved BGM. Define the canonical marker — `music: none` in the storyboard's top YAML block, plus no SCRIPT.md — and honor it: audio generate produces nothing (removing stale audio_meta.json, since absence is what assemble treats as silent), and `music: none` with narration keeps TTS while turning BGM off. Also documents the `auth status` exit-code contract (exit 1 while signed out is the normal offline state, not a failure) in the product-launch Step 0 note and the CLI skill's cloud reference. * fix(skills): transient-init retry for standalone animation-map and contrast-report The standalone helpers called initializeSession exactly once, so a valid modular project — whose sub-composition timelines register asynchronously — could hit the readiness deadline and die with the transient "zero duration / Runtime ready: false" diagnostic the render pipeline retries (probeStage). Add initializeSessionWithRetry to the shared package-loader (both byte-identical copies): close the crashed session and retry once with a fresh browser, gated by the engine's canonical isTransientBrowserError — now re-exported from @hyperframes/producer, with a frozen fallback pattern list for older published packages. The "Runtime ready: true" fast-fail (a genuine authoring bug) still fails without a retry. * feat(skills): extend the fully-silent marker to faceless-explainer and pr-to-video Both workflows reuse product-launch's audio model — their Step 3.1 gates carried the same undefined "marked silent" phrase, and their (intentionally identical) audio.mjs copies had the same unconditional BGM retrieve. Port the `music: none` marker handling into both copies, define the marker in their SKILL.md Step 3.1 and story-design references, and turn the copies' "intentionally identical" header claim into a byte-identity pin test so the next fix can't silently miss one of them. * test(cli): reset the prune mock explicitly instead of relying on restoreAllMocks The converge test's toHaveBeenCalledTimes(1) held only because vitest 3's vi.restoreAllMocks() clears vi.fn() call state; vitest 4 restores spies only, so the count would accumulate across tests and fail. Reset pruneOrphanedLockEntries in beforeEach like the other manifest mocks — passes under both vitest 3.2.4 (pinned) and vitest 4. * test(skills): close review findings — package-loader pin, whisper win32 parity, quoted-none Review follow-ups on #2476: - package-loader.mjs byte-identity pin (the elevated concern): the two copies now carry initializeSessionWithRetry + FALLBACK_TRANSIENT_PATTERNS, exactly the shared-logic shape a future fix could land in one copy and miss in the other — same enforcement as the audio.mjs pin. - whisper win32 call-site parity: runWhisper's npx resolution lifted into lib/npx-sync.mjs (resolveNpxInvocation, injectable params matching the localTtsGenerate idiom) with the same three-branch coverage as the Kokoro site — plus the hard-fail contract (throws actionably, since the whisper fallback has no next provider to fall through to). - quoted music: "none" pin: the vendored storyboard parser strips matching quotes at parse time (stripQuotes), so the silent marker already accepts the quoted spelling — pinned so that stays true.
93 lines
3.7 KiB
JavaScript
93 lines
3.7 KiB
JavaScript
import { execFileSync } from "node:child_process";
|
|
import { existsSync, statSync } from "node:fs";
|
|
import { tmpdir } from "node:os";
|
|
import { join } from "node:path";
|
|
import { resolveSpawnCommand } from "../../audio/scripts/lib/tts.mjs";
|
|
|
|
// Local voiceover via the packaged Kokoro-82M TTS (the `hyperframes tts` CLI),
|
|
// the free/private default now that HeyGen TTS costs wallet credits. Kokoro runs
|
|
// on-device (CPU, faster-than-realtime, bundled voices, native word timestamps),
|
|
// so no key and no per-call charge. When Kokoro is not set up, this returns null
|
|
// and the registry falls through to the HeyGen TTS upsell.
|
|
//
|
|
// Delegated to the hyperframes CLI (same as transcribe / remove-background), not
|
|
// re-implemented here. ffprobe reads the duration back for the ledger.
|
|
|
|
function probeDurationSeconds(file) {
|
|
try {
|
|
const out = execFileSync(
|
|
"ffprobe",
|
|
["-v", "error", "-show_entries", "format=duration", "-of", "csv=p=0", file],
|
|
{ encoding: "utf8", timeout: 15000 },
|
|
);
|
|
const d = parseFloat(String(out).trim());
|
|
return Number.isFinite(d) ? d : undefined;
|
|
} catch {
|
|
return undefined;
|
|
}
|
|
}
|
|
|
|
// `platform`/`execFn`/`env`/`pathExists` params (defaulting to the real
|
|
// values) exist so tests can exercise the win32 branch without mocking
|
|
// node:child_process (its ESM exports are non-configurable) — same idiom as
|
|
// spawnP in ../../audio/scripts/lib/tts.mjs.
|
|
export async function localTtsGenerate(
|
|
intent,
|
|
ctx,
|
|
platform = process.platform,
|
|
execFn = execFileSync,
|
|
env = process.env,
|
|
pathExists = existsSync,
|
|
) {
|
|
const outPath = join(tmpdir(), `media-use-kokoro-${process.pid}-${Date.now()}.wav`);
|
|
const argv = ["hyperframes", "tts", intent, "--output", outPath];
|
|
if (ctx?.voice) argv.push("--voice", ctx.voice);
|
|
if (ctx?.lang && ctx.lang !== "en") argv.push("--lang", ctx.lang);
|
|
// On Windows a bare "npx" is npx.cmd, which execFileSync cannot exec
|
|
// (spawnSync npx ENOENT) — resolveSpawnCommand reroutes it through
|
|
// node + npx-cli.js, same as the audio engine's TTS spawns.
|
|
const resolved = resolveSpawnCommand(
|
|
"npx",
|
|
argv,
|
|
{ encoding: "utf8", timeout: 300000, stdio: ["ignore", "pipe", "pipe"] },
|
|
platform,
|
|
env,
|
|
pathExists,
|
|
);
|
|
if (!resolved) {
|
|
// npx-on-win32 with no resolvable npx-cli.js — same terminal condition
|
|
// spawnP warns about. Fall through to the next provider rather than crash.
|
|
console.error(
|
|
"media-use: local voice not enabled (kokoro). Cannot run npx on Windows: " +
|
|
"npm's npx-cli.js was not found (install npm with Node, or run via npx/npm run so npm_execpath is set).",
|
|
);
|
|
return null;
|
|
}
|
|
try {
|
|
execFn(resolved.cmd, resolved.args, resolved.opts);
|
|
} catch (err) {
|
|
// `hyperframes tts` prints its "kokoro-onnx not installed" hint to stdout
|
|
// (clack UI), so read both streams and surface the actionable enable-command
|
|
// rather than a bare "Command failed": otherwise resolve silently falls
|
|
// through to the PAID HeyGen TTS upsell when free local voice was one pip away.
|
|
const out = `${err.stdout?.toString() ?? ""}${err.stderr?.toString() ?? ""}`.trim();
|
|
const hint = /not installed|pip install kokoro/i.test(out)
|
|
? "install for free on-device voice: pip install kokoro-onnx soundfile (or set HYPERFRAMES_PYTHON to a venv that has it)"
|
|
: out.slice(-200) || err.message;
|
|
console.error(`media-use: local voice not enabled (kokoro). ${hint}`);
|
|
return null;
|
|
}
|
|
if (!existsSync(outPath) || statSync(outPath).size === 0) return null;
|
|
return {
|
|
localPath: outPath,
|
|
ext: ".wav",
|
|
source: "generated",
|
|
metadata: {
|
|
description: intent,
|
|
provider: "kokoro.local",
|
|
duration: probeDurationSeconds(outPath),
|
|
provenance: { engine: "kokoro-82m", prompt: intent },
|
|
},
|
|
};
|
|
}
|