mirror of
https://github.com/heygen-com/hyperframes.git
synced 2026-09-04 07:19:52 +00:00
* fix(cli): invalidate the skills nudge cache after a successful install/update/check The passive "N skills out of date or missing" nudge reads a 24h config cache that only the background check (on non-skills commands) ever wrote. The skills commands themselves are excluded from the nudge pipeline, so a successful `skills update`/install/check never refreshed or dropped the cached verdict — the pre-install count kept printing on every other command for up to 24h. Reconcile commands now drop the cached verdict (counts + timestamp) so the next command's background check re-runs for real. The offline presence-only path deliberately keeps the cache: that run learned nothing about freshness. * fix(skills): win32-safe npx spawns in media-use + accurate whisper wording The Whisper transcribe fallback and the Kokoro local-TTS delegation both spawned a bare "npx" via execFileSync — on Windows npx is npx.cmd, which spawn cannot exec, so both paths died with `spawnSync npx ENOENT`. Route them through the skill's existing resolveSpawnCommand (node + npx-cli.js on win32, no shell:true), same as the audio engine's TTS spawns. Also corrects the "bundled with the hyperframes CLI" claim about whisper.cpp: it is resolved from PATH / installed via Homebrew / built from source with git+cmake on first use, and models download from HuggingFace — nothing whisper is shipped in the package. * feat(skills): canonical fully-silent marker + auth status exit-code docs product-launch's Step 3.1 gate said "or the project is marked silent" but nothing defined how to mark one, and audio.mjs unconditionally retrieved BGM. Define the canonical marker — `music: none` in the storyboard's top YAML block, plus no SCRIPT.md — and honor it: audio generate produces nothing (removing stale audio_meta.json, since absence is what assemble treats as silent), and `music: none` with narration keeps TTS while turning BGM off. Also documents the `auth status` exit-code contract (exit 1 while signed out is the normal offline state, not a failure) in the product-launch Step 0 note and the CLI skill's cloud reference. * fix(skills): transient-init retry for standalone animation-map and contrast-report The standalone helpers called initializeSession exactly once, so a valid modular project — whose sub-composition timelines register asynchronously — could hit the readiness deadline and die with the transient "zero duration / Runtime ready: false" diagnostic the render pipeline retries (probeStage). Add initializeSessionWithRetry to the shared package-loader (both byte-identical copies): close the crashed session and retry once with a fresh browser, gated by the engine's canonical isTransientBrowserError — now re-exported from @hyperframes/producer, with a frozen fallback pattern list for older published packages. The "Runtime ready: true" fast-fail (a genuine authoring bug) still fails without a retry. * feat(skills): extend the fully-silent marker to faceless-explainer and pr-to-video Both workflows reuse product-launch's audio model — their Step 3.1 gates carried the same undefined "marked silent" phrase, and their (intentionally identical) audio.mjs copies had the same unconditional BGM retrieve. Port the `music: none` marker handling into both copies, define the marker in their SKILL.md Step 3.1 and story-design references, and turn the copies' "intentionally identical" header claim into a byte-identity pin test so the next fix can't silently miss one of them. * test(cli): reset the prune mock explicitly instead of relying on restoreAllMocks The converge test's toHaveBeenCalledTimes(1) held only because vitest 3's vi.restoreAllMocks() clears vi.fn() call state; vitest 4 restores spies only, so the count would accumulate across tests and fail. Reset pruneOrphanedLockEntries in beforeEach like the other manifest mocks — passes under both vitest 3.2.4 (pinned) and vitest 4. * test(skills): close review findings — package-loader pin, whisper win32 parity, quoted-none Review follow-ups on #2476: - package-loader.mjs byte-identity pin (the elevated concern): the two copies now carry initializeSessionWithRetry + FALLBACK_TRANSIENT_PATTERNS, exactly the shared-logic shape a future fix could land in one copy and miss in the other — same enforcement as the audio.mjs pin. - whisper win32 call-site parity: runWhisper's npx resolution lifted into lib/npx-sync.mjs (resolveNpxInvocation, injectable params matching the localTtsGenerate idiom) with the same three-branch coverage as the Kokoro site — plus the hard-fail contract (throws actionably, since the whisper fallback has no next provider to fall through to). - quoted music: "none" pin: the vendored storyboard parser strips matching quotes at parse time (stripQuotes), so the silent marker already accepts the quoted spelling — pinned so that stays true.
181 lines
6.5 KiB
JavaScript
181 lines
6.5 KiB
JavaScript
#!/usr/bin/env node
|
|
|
|
import { execFileSync } from "node:child_process";
|
|
import {
|
|
existsSync,
|
|
readFileSync,
|
|
writeFileSync,
|
|
copyFileSync,
|
|
renameSync,
|
|
mkdtempSync,
|
|
rmSync,
|
|
} from "node:fs";
|
|
import { homedir, tmpdir } from "node:os";
|
|
import { basename, extname, join, resolve } from "node:path";
|
|
import { parseArgs } from "node:util";
|
|
import { mergeTokensToWords } from "./lib/parakeet-words.mjs";
|
|
import { track } from "./lib/telemetry.mjs";
|
|
import { resolveNpxInvocation } from "./lib/npx-sync.mjs";
|
|
|
|
// The DEFAULT local transcription path. Prefers NVIDIA Parakeet-TDT via
|
|
// parakeet-mlx, which beats whisper.cpp on the Open ASR Leaderboard (~6.05% vs
|
|
// 7.44% avg WER, and 4.73% vs 5.96% on noisy test-other) and is 5-10x faster
|
|
// with native punctuation. Emits { text, words:[{text,start,end}] } (word
|
|
// timestamps merged from Parakeet's sub-word tokens) for transcript-cut /
|
|
// captions / the audio engine.
|
|
//
|
|
// Parakeet v3 covers English + 25 European languages. For other languages, or
|
|
// when parakeet-mlx is not installed, it falls back to whisper.cpp via
|
|
// `hyperframes transcribe` (99 languages; the CLI resolves/builds whisper.cpp
|
|
// on first use — it is not bundled). `--engine` forces one.
|
|
|
|
const { values: args } = parseArgs({
|
|
options: {
|
|
input: { type: "string", short: "i" },
|
|
out: { type: "string", short: "o" },
|
|
engine: { type: "string", default: "auto" }, // auto | parakeet | whisper
|
|
model: { type: "string", default: "mlx-community/parakeet-tdt-0.6b-v3" },
|
|
json: { type: "boolean", default: false },
|
|
help: { type: "boolean", short: "h", default: false },
|
|
},
|
|
strict: true,
|
|
});
|
|
|
|
if (args.help) {
|
|
console.log(`media-use transcribe: better-than-whisper local ASR (Parakeet), whisper.cpp fallback
|
|
|
|
Usage:
|
|
node transcribe.mjs --input audio.wav [--out audio.transcribe.json] [--engine auto|parakeet|whisper]
|
|
|
|
Parakeet (default) beats whisper.cpp on accuracy + speed for English/European
|
|
languages; whisper.cpp (99 languages) is the fallback. Install Parakeet once:
|
|
uv venv ~/.venvs/parakeet && VIRTUAL_ENV=~/.venvs/parakeet uv pip install parakeet-mlx`);
|
|
process.exit(0);
|
|
}
|
|
|
|
if (!args.input) {
|
|
console.error("error: --input is required");
|
|
process.exit(2);
|
|
}
|
|
const inputPath = resolve(args.input);
|
|
if (!existsSync(inputPath)) {
|
|
console.error(`error: input not found: ${inputPath}`);
|
|
process.exit(2);
|
|
}
|
|
const outPath = resolve(
|
|
args.out || `${inputPath.slice(0, -extname(inputPath).length)}.transcribe.json`,
|
|
);
|
|
|
|
// Locate the parakeet-mlx runner the same way the CLI does: env override, then
|
|
// the documented ~/.venvs/parakeet install, then PATH. Checking the venv (not
|
|
// just PATH) is what keeps a user who followed the install docs verbatim from
|
|
// silently falling through to whisper. Returns the runner path, or null.
|
|
function resolveParakeet() {
|
|
for (const p of [
|
|
process.env.HYPERFRAMES_PARAKEET,
|
|
join(homedir(), ".venvs", "parakeet", "bin", "parakeet-mlx"),
|
|
]) {
|
|
if (p && existsSync(p)) return p;
|
|
}
|
|
try {
|
|
execFileSync("parakeet-mlx", ["--help"], {
|
|
stdio: ["ignore", "ignore", "ignore"],
|
|
timeout: 20000,
|
|
});
|
|
return "parakeet-mlx";
|
|
} catch {
|
|
return null;
|
|
}
|
|
}
|
|
|
|
// Write via a sibling temp + atomic rename so a SIGKILL mid-write can't leave a
|
|
// truncated transcript at outPath (downstream reads it as valid JSON).
|
|
function atomicWrite(target, data) {
|
|
const tmp = `${target}.tmp-${process.pid}`;
|
|
writeFileSync(tmp, data);
|
|
renameSync(tmp, target);
|
|
}
|
|
|
|
function report(engine, wordCount) {
|
|
if (args.json) console.log(JSON.stringify({ ok: true, out: outPath, engine, words: wordCount }));
|
|
else
|
|
console.log(
|
|
`transcribed ${basename(inputPath)} -> ${outPath}${wordCount != null ? ` (${wordCount} words,` : " ("}${engine})`,
|
|
);
|
|
}
|
|
|
|
function runParakeet(runner) {
|
|
const workDir = mkdtempSync(join(tmpdir(), "media-use-asr-"));
|
|
try {
|
|
execFileSync(
|
|
runner,
|
|
[inputPath, "--model", args.model, "--output-format", "json", "--output-dir", workDir],
|
|
{ stdio: ["ignore", "pipe", "pipe"], timeout: 1_800_000 },
|
|
);
|
|
const jsonPath = join(workDir, `${basename(inputPath, extname(inputPath))}.json`);
|
|
if (!existsSync(jsonPath)) throw new Error("parakeet produced no JSON");
|
|
const merged = mergeTokensToWords(JSON.parse(readFileSync(jsonPath, "utf8")));
|
|
atomicWrite(outPath, JSON.stringify(merged, null, 2));
|
|
report("parakeet", merged.words.length);
|
|
} finally {
|
|
rmSync(workDir, { recursive: true, force: true });
|
|
}
|
|
}
|
|
|
|
// whisper.cpp via the hyperframes CLI (fetched/built on first use — see
|
|
// SKILL.md): writes transcript.json into --dir; relocate to --out.
|
|
function runWhisper() {
|
|
const workDir = mkdtempSync(join(tmpdir(), "media-use-whisper-"));
|
|
try {
|
|
// On Windows a bare "npx" is npx.cmd, which execFileSync cannot exec
|
|
// (spawnSync npx ENOENT) — resolveNpxInvocation reroutes it through
|
|
// node + npx-cli.js (and throws actionably when it can't), same
|
|
// mechanism as the audio engine's TTS spawns.
|
|
const resolved = resolveNpxInvocation(
|
|
["hyperframes", "transcribe", inputPath, "--dir", workDir],
|
|
{ stdio: ["ignore", "pipe", "pipe"], timeout: 1_800_000 },
|
|
);
|
|
execFileSync(resolved.cmd, resolved.args, resolved.opts);
|
|
const produced = join(workDir, "transcript.json");
|
|
if (!existsSync(produced)) throw new Error("whisper produced no transcript.json");
|
|
const tmp = `${outPath}.tmp-${process.pid}`;
|
|
copyFileSync(produced, tmp);
|
|
renameSync(tmp, outPath); // atomic publish
|
|
let words;
|
|
try {
|
|
const t = JSON.parse(readFileSync(outPath, "utf8"));
|
|
words = Array.isArray(t?.words) ? t.words.length : undefined;
|
|
} catch {
|
|
/* leave undefined */
|
|
}
|
|
report("whisper", words);
|
|
} finally {
|
|
rmSync(workDir, { recursive: true, force: true });
|
|
}
|
|
}
|
|
|
|
try {
|
|
const parakeetBin = resolveParakeet();
|
|
const engine =
|
|
args.engine === "parakeet" || args.engine === "whisper"
|
|
? args.engine
|
|
: parakeetBin
|
|
? "parakeet"
|
|
: "whisper";
|
|
if (engine === "parakeet") {
|
|
if (!parakeetBin) {
|
|
throw new Error(
|
|
"parakeet-mlx not found (checked $HYPERFRAMES_PARAKEET, ~/.venvs/parakeet, and PATH). Install: uv venv ~/.venvs/parakeet && VIRTUAL_ENV=~/.venvs/parakeet uv pip install parakeet-mlx (or use --engine whisper)",
|
|
);
|
|
}
|
|
runParakeet(parakeetBin);
|
|
} else {
|
|
runWhisper();
|
|
}
|
|
await track("media_use_transcribe", { engine });
|
|
} catch (err) {
|
|
if (args.json) console.log(JSON.stringify({ ok: false, error: err.message }));
|
|
else console.error(`error: transcription failed: ${err.message}`);
|
|
process.exit(1);
|
|
}
|