mirror of
https://github.com/heygen-com/hyperframes.git
synced 2026-09-05 10:14:30 +00:00
* fix: handle caption skin workflow * docs(skills): simplify the finalize step across video workflows - Drop --strict-layout; all skills use plain `hyperframes inspect` - Add the caption text_box_overflow false-positive note to faceless-explainer - On a failed check, the orchestrator makes the cheapest safe edit itself (no worker re-dispatch / Step 3 backtrack language) - Snapshot: glance at the stitched contact-sheet.jpg and move on Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(auth): onboarding-first `auth status` + shared TTS/BGM preflight When no HeyGen credential is configured, `hyperframes auth status` now prints registration-first guidance instead of a terse error: - Interactive / agent-driven sessions get sign-in guidance led by `hyperframes auth login` (the OAuth step that also creates an account and is shared with heygen-cli), and never steer users to a per-repo `.env`. CI / non-interactive runs get a terse note. Exit 1 is kept so the "am I logged in?" `$?` contract still holds. - It probes which local engine voice/music will fall back to (Kokoro / MusicGen, mirroring the skill resolution order) and whether their Python deps are installed, with a pip hint when missing. `--json` exposes `recommended_action` + `offline_engines` for skills to branch. - `doctor` gains matching "TTS (Kokoro)" / "BGM (MusicGen)" checks via the same shared probe (findPython/hasPythonModules extracted to tts/python.ts; provider resolution in audio/providers.ts). Every TTS/BGM workflow now relays this at Step 0 (setup) instead of improvising its own "missing key" prompt: pr-to-video, product-launch- video, faceless-explainer, website-to-video, music-to-video. The canonical behavior + key-priority table live once in hyperframes-media. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(pr-to-video): scale recommended video length to PR change size Step 0 led with a fixed ~60-90s length default. Now the recommended length is derived from the PR's diff stat (lines added+deleted, nudged by file count) on a tier scale (trivial ~20-40s → large ~110-180s, hard cap ~3 min), reusing the same PR peek already done to infer the angle. The agent states the basis when proposing it, and a huge PR with one headline change still stays tight. User can always override. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(captions): embed brand fonts whose files use separators brandFontFaces() matched font files by stripping only whitespace, so an underscore/hyphen-named file (TT_Norms_Pro_Bold.woff2) never matched the family key "ttnormspro" — captions shipped with no @font-face, the font_family_without_font_face bug. Now both family and filename normalize away all non-alphanumerics; families match longest-key-first so a parent family can't swallow a more specific one's files (TT Norms Pro vs Mono); each file is claimed once; "demibold" ranks before "bold"; and when nothing matches it warns loudly at build time instead of returning "". Also: parseFonts() falls back to h1/h2/title/hero display roles, and the frame-worker + caption authoring docs spell out that only shipped font files render — no system CJK/Devanagari families on the headless renderer. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(hyperframes-media): enforce sign-in preflight on standalone BGM/TTS A one-off "generate me a BGM" request went straight to local MusicGen without recommending sign-in: bgm.md/tts.md framed the no-credential path as an automatic fallback, so the generation path bypassed the Preflight stop, and the preflight used a bare `hyperframes auth status` that isn't on PATH in a fresh `npx skills` project. - Preflight now applies to one-off generation as well as workflows, uses `npx hyperframes auth status`, and says: if the CLI can't run, still recommend signing in and STOP — never treat "no credential" as a silent green light for local generation. - bgm.md and tts.md point at the Preflight before generating, reframing local generation as the fallback the user opts into, not a default. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(auth): add Authentication & API keys guide Document signing in, the keys each capability (voice, music, capture) uses, their resolution priority, and the fully local fallback. Add the guide to the nav and cross-link it from the cloud deploy note and the CLI env-var reference. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(lint): strip HTML comments in a fixpoint loop (CodeQL) Single-pass <!-- --> removal can re-form a complete comment from adjacent markers (e.g. `<<!-- -->!-- ... -->`), letting a decoy <template> survive and hijack the template-boundary match. Loop to a fixpoint, mirroring the captions.mjs precedent; add a regression test that fails on single-pass (2 root findings) and passes on the loop. Also wrap the build-frame.mjs node:fs imports to satisfy oxfmt — the new copyFileSync import pushed the line past the width limit, which was the sole cause of the Format / Preflight CI failures. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(lint): strip HTML comments with a linear scan (CodeQL ReDoS) The fixpoint loop still ran a /<!--[\s\S]*?-->/ regex per pass, which backtracks O(n^2) on inputs with many unterminated "<!--" — CodeQL js/polynomial-redos (high). Looping the same regex (the prescribed fix) never addressed this; only the regex itself does. Replace it with an indexOf-based linear strip in utils.ts (stripHtmlComments), kept in a fixpoint loop so markers that re-form when a comment is removed are still stripped. 200k unterminated "<!--" now strips in ~3ms instead of quadratic time; behavior is otherwise unchanged — unterminated comments are kept verbatim, as the old regex left them. The re-forming regression test still guards it. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(auth): make TTS/BGM sign-in guidance accurate and runnable From team review of the not-signed-in onboarding: - OAuth is a `hyperframes auth login` feature only. The separate `heygen` CLI is API-key-only — `heygen auth login` stores a pasted key, it is not OAuth and does not create an account. Stop presenting the two CLIs as the same OAuth/sign-up step. - Use `npx hyperframes` in every imperative and runtime hint. Bare `hyperframes` is not on PATH on a fresh machine (command not found); only `npx hyperframes` is guaranteed. Also updates the JSON recommended_action. - Drop `heygen auth login` from the terminal/skill onboarding: it needs its own install and there is no `npx heygen`, so it was a command-not-found trap. The shared-credential fact stays in the reference docs. Covers the `auth status` guidance + tests, the Authentication docs, the shared hyperframes-media preflight (SKILL, requirements, tts, error hints), and the `npx hyperframes auth status` preflight in every TTS/BGM workflow (pr-to-video, product-launch-video, faceless-explainer, website-to-video, music-to-video). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
217 lines
9.3 KiB
JavaScript
217 lines
9.3 KiB
JavaScript
// tts.mjs — multi-provider TTS for the media audio engine. The provider chain,
|
|
// auto-detected from env, is the one documented in ../SKILL.md:
|
|
//
|
|
// 1. HeyGen (Starfish) — $HEYGEN_API_KEY / $HYPERFRAMES_API_KEY / ~/.heygen.
|
|
// Direct v3 REST (NOT `hyperframes tts`, which in the published build is
|
|
// Kokoro-only and silently ignores a HeyGen key). Returns word_timestamps
|
|
// in the same call, so no separate transcribe pass.
|
|
// 2. ElevenLabs — $ELEVENLABS_API_KEY + `pip install elevenlabs`. No
|
|
// word timings → caller chains transcribeWav().
|
|
// 3. Kokoro-82M (local) — always available, via the published `hyperframes tts`
|
|
// CLI. No word timings → caller chains transcribeWav().
|
|
//
|
|
// "HeyGen available" is decided by CREDENTIAL presence (heygenCredential), never
|
|
// by the CLI — see the note above.
|
|
|
|
import { spawn, spawnSync } from "node:child_process";
|
|
import { existsSync, mkdirSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs";
|
|
import { tmpdir } from "node:os";
|
|
import { dirname, join } from "node:path";
|
|
import { heygenAuthHeaders, heygenCredential, heygenJSON } from "./heygen.mjs";
|
|
|
|
// ── provider detection ────────────────────────────────────────────────────────
|
|
export function heygenAvailable() {
|
|
return heygenCredential() !== null;
|
|
}
|
|
export function elevenlabsAvailable() {
|
|
if (!process.env.ELEVENLABS_API_KEY) return false;
|
|
const r = spawnSync("python3", ["-c", "import elevenlabs"], { stdio: "ignore" });
|
|
return r.status === 0;
|
|
}
|
|
|
|
// First available provider wins; an explicit choice is honored (and validated).
|
|
export function pickProvider(userProvider) {
|
|
if (userProvider) {
|
|
if (!["heygen", "elevenlabs", "kokoro"].includes(userProvider))
|
|
throw new Error(`invalid provider "${userProvider}" (heygen | elevenlabs | kokoro)`);
|
|
if (userProvider === "heygen" && !heygenAvailable())
|
|
throw new Error(
|
|
"provider=heygen but no HeyGen credentials (set $HEYGEN_API_KEY or run `npx hyperframes auth login`)",
|
|
);
|
|
if (userProvider === "elevenlabs" && !process.env.ELEVENLABS_API_KEY)
|
|
throw new Error("provider=elevenlabs but $ELEVENLABS_API_KEY is not set");
|
|
return userProvider;
|
|
}
|
|
return heygenAvailable() ? "heygen" : elevenlabsAvailable() ? "elevenlabs" : "kokoro";
|
|
}
|
|
|
|
// ── voice resolution ──────────────────────────────────────────────────────────
|
|
// HeyGen /v3/voices/speech only accepts STARFISH voice_ids; auto-pick the first
|
|
// English public starfish voice when none is pinned. ElevenLabs/Kokoro have
|
|
// their own defaults.
|
|
export async function resolveVoiceId({ provider, userVoice, lang = "en" }) {
|
|
if (userVoice) return userVoice;
|
|
if (provider === "elevenlabs") return "21m00Tcm4TlvDq8ikWAM"; // Rachel
|
|
if (provider === "kokoro") {
|
|
if (lang === "en") return "am_michael";
|
|
throw new Error("Kokoro non-English needs an explicit --voice (see references/tts.md)");
|
|
}
|
|
// heygen
|
|
const payload = await heygenJSON(`/voices?engine=starfish&type=public&limit=50`, {
|
|
headers: heygenAuthHeaders(),
|
|
});
|
|
const voices = payload.data ?? payload.voices ?? [];
|
|
const pick = voices.find((v) => v.language === "English") ?? voices[0];
|
|
if (!pick) throw new Error("no public starfish voice to default to — pass --voice");
|
|
return pick.voice_id;
|
|
}
|
|
|
|
// ── helpers ─────────────────────────────────────────────────────────────────
|
|
export function withWordIds(words) {
|
|
return (words ?? []).map((w, i) => ({ id: `w${i}`, text: w.text, start: w.start, end: w.end }));
|
|
}
|
|
|
|
export function ffprobeDuration(absPath) {
|
|
const r = spawnSync(
|
|
"ffprobe",
|
|
["-v", "error", "-show_entries", "format=duration", "-of", "default=nw=1:nk=1", absPath],
|
|
{ encoding: "utf8" },
|
|
);
|
|
if (r.status !== 0) return NaN;
|
|
return parseFloat(String(r.stdout).trim());
|
|
}
|
|
|
|
function spawnP(cmd, args, opts) {
|
|
return new Promise((resolve) => {
|
|
const p = spawn(cmd, args, { stdio: "ignore", ...opts });
|
|
p.on("exit", (code) => resolve({ status: code ?? -1 }));
|
|
p.on("error", () => resolve({ status: -1 }));
|
|
});
|
|
}
|
|
|
|
// mp3/whatever bytes → wav 44.1k mono at destWav (ffmpeg detects true format).
|
|
function transcodeToWav(bytes, destWav) {
|
|
const td = mkdtempSync(join(tmpdir(), "hf-tts-"));
|
|
const tmp = join(td, "a.mp3");
|
|
writeFileSync(tmp, bytes);
|
|
mkdirSync(dirname(destWav), { recursive: true });
|
|
const ff = spawnSync(
|
|
"ffmpeg",
|
|
["-y", "-loglevel", "error", "-i", tmp, "-ar", "44100", "-ac", "1", destWav],
|
|
{ stdio: "ignore" },
|
|
);
|
|
rmSync(td, { recursive: true, force: true });
|
|
return ff.status === 0 && existsSync(destWav);
|
|
}
|
|
|
|
const ELEVENLABS_PY = `
|
|
import os, sys
|
|
from elevenlabs.client import ElevenLabs
|
|
from elevenlabs import save
|
|
client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
|
|
text = open(sys.argv[1]).read()
|
|
audio = client.text_to_speech.convert(
|
|
text=text, voice_id=sys.argv[2],
|
|
model_id="eleven_multilingual_v2", output_format="mp3_44100_128",
|
|
)
|
|
save(audio, sys.argv[3])
|
|
`;
|
|
|
|
// ── synthesize one line ───────────────────────────────────────────────────────
|
|
// Writes wav at wavAbs. Returns { ok, words } — words is the raw
|
|
// [{text,start,end}] array for HeyGen (native), or null for ElevenLabs/Kokoro
|
|
// (caller must transcribeWav). Never throws; failures return { ok:false }.
|
|
export async function synthesizeOne({
|
|
provider,
|
|
text,
|
|
voiceId,
|
|
lang = "en",
|
|
speed = 1.0,
|
|
wavAbs,
|
|
hyperframesDir,
|
|
}) {
|
|
if (provider === "heygen") return synthesizeHeygen({ text, voiceId, lang, speed, wavAbs });
|
|
if (provider === "elevenlabs") {
|
|
const r = await spawnP(
|
|
"python3",
|
|
["-c", ELEVENLABS_PY, writeTmpText(text), voiceId, wavAbs],
|
|
{},
|
|
);
|
|
return { ok: r.status === 0 && existsSync(wavAbs), words: null };
|
|
}
|
|
// kokoro — via the published CLI; --output is relative to the project dir.
|
|
const wavRel = relTo(hyperframesDir, wavAbs);
|
|
const args = ["hyperframes", "tts", writeTmpText(text), "--voice", voiceId, "--output", wavRel];
|
|
if (lang !== "en") args.push("--lang", lang);
|
|
const r = await spawnP("npx", args, { cwd: hyperframesDir });
|
|
return { ok: r.status === 0 && existsSync(wavAbs), words: null };
|
|
}
|
|
|
|
async function synthesizeHeygen({ text, voiceId, lang, speed, wavAbs }) {
|
|
try {
|
|
const body = { text, voice_id: voiceId, speed };
|
|
if (lang !== "en") body.language = lang;
|
|
const payload = await heygenJSON(`/voices/speech`, {
|
|
method: "POST",
|
|
headers: heygenAuthHeaders(),
|
|
body,
|
|
});
|
|
const inner = payload.data ?? payload;
|
|
if (!inner.audio_url) return { ok: false, words: null };
|
|
const res = await fetch(inner.audio_url);
|
|
if (!res.ok) return { ok: false, words: null };
|
|
const bytes = Buffer.from(await res.arrayBuffer());
|
|
// .wav output → transcode to 44.1k mono; .mp3 → raw bytes (no ffmpeg). The
|
|
// engine always asks for .wav; the standalone heygen-tts CLI may ask for .mp3.
|
|
if (wavAbs.endsWith(".wav")) {
|
|
if (!transcodeToWav(bytes, wavAbs)) return { ok: false, words: null };
|
|
} else {
|
|
mkdirSync(dirname(wavAbs), { recursive: true });
|
|
writeFileSync(wavAbs, bytes);
|
|
}
|
|
const words = Array.isArray(inner.word_timestamps)
|
|
? inner.word_timestamps
|
|
.filter((w) => w && typeof w.word === "string" && isFinite(w.start) && isFinite(w.end))
|
|
.filter((w) => !/^<.*>$/.test(w.word.trim())) // drop <start>/<end> sentinels
|
|
.map((w) => ({ text: w.word, start: w.start, end: w.end }))
|
|
: [];
|
|
return { ok: true, words };
|
|
} catch {
|
|
return { ok: false, words: null };
|
|
}
|
|
}
|
|
|
|
// ElevenLabs/Kokoro have no word timings — run Whisper over the wav. Returns the
|
|
// flat [{id,text,start,end}] word array, or null. Each call uses a throwaway
|
|
// --dir so parallel scenes don't collide on transcript.json.
|
|
export async function transcribeWav({ wavRel, lang = "en", hyperframesDir }) {
|
|
const model = lang === "en" ? "small.en" : "small";
|
|
const td = mkdtempSync(join(tmpdir(), "hf-trans-"));
|
|
const args = ["hyperframes", "transcribe", wavRel, "--model", model, "--dir", td];
|
|
if (lang !== "en") args.push("--language", lang);
|
|
const r = await spawnP("npx", args, { cwd: hyperframesDir });
|
|
let words = null;
|
|
if (r.status === 0) {
|
|
const src = join(td, "transcript.json");
|
|
if (existsSync(src)) {
|
|
try {
|
|
const arr = JSON.parse(readFileSync(src, "utf8"));
|
|
if (Array.isArray(arr) && arr.length) words = arr;
|
|
} catch {}
|
|
}
|
|
}
|
|
rmSync(td, { recursive: true, force: true });
|
|
return words;
|
|
}
|
|
|
|
// ── tiny local utils ──────────────────────────────────────────────────────────
|
|
function writeTmpText(text) {
|
|
const td = mkdtempSync(join(tmpdir(), "hf-txt-"));
|
|
const p = join(td, "line.txt");
|
|
writeFileSync(p, text);
|
|
return p;
|
|
}
|
|
function relTo(base, abs) {
|
|
return abs.startsWith(base + "/") ? abs.slice(base.length + 1) : abs;
|
|
}
|