mirror of
https://github.com/heygen-com/hyperframes.git
synced 2026-09-10 22:20:14 +00:00
feat(cli): shared TTS/BGM auth preflight + caption and skill-workflow fixes (#1697)
* fix: handle caption skin workflow * docs(skills): simplify the finalize step across video workflows - Drop --strict-layout; all skills use plain `hyperframes inspect` - Add the caption text_box_overflow false-positive note to faceless-explainer - On a failed check, the orchestrator makes the cheapest safe edit itself (no worker re-dispatch / Step 3 backtrack language) - Snapshot: glance at the stitched contact-sheet.jpg and move on Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(auth): onboarding-first `auth status` + shared TTS/BGM preflight When no HeyGen credential is configured, `hyperframes auth status` now prints registration-first guidance instead of a terse error: - Interactive / agent-driven sessions get sign-in guidance led by `hyperframes auth login` (the OAuth step that also creates an account and is shared with heygen-cli), and never steer users to a per-repo `.env`. CI / non-interactive runs get a terse note. Exit 1 is kept so the "am I logged in?" `$?` contract still holds. - It probes which local engine voice/music will fall back to (Kokoro / MusicGen, mirroring the skill resolution order) and whether their Python deps are installed, with a pip hint when missing. `--json` exposes `recommended_action` + `offline_engines` for skills to branch. - `doctor` gains matching "TTS (Kokoro)" / "BGM (MusicGen)" checks via the same shared probe (findPython/hasPythonModules extracted to tts/python.ts; provider resolution in audio/providers.ts). Every TTS/BGM workflow now relays this at Step 0 (setup) instead of improvising its own "missing key" prompt: pr-to-video, product-launch- video, faceless-explainer, website-to-video, music-to-video. The canonical behavior + key-priority table live once in hyperframes-media. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(pr-to-video): scale recommended video length to PR change size Step 0 led with a fixed ~60-90s length default. Now the recommended length is derived from the PR's diff stat (lines added+deleted, nudged by file count) on a tier scale (trivial ~20-40s → large ~110-180s, hard cap ~3 min), reusing the same PR peek already done to infer the angle. The agent states the basis when proposing it, and a huge PR with one headline change still stays tight. User can always override. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(captions): embed brand fonts whose files use separators brandFontFaces() matched font files by stripping only whitespace, so an underscore/hyphen-named file (TT_Norms_Pro_Bold.woff2) never matched the family key "ttnormspro" — captions shipped with no @font-face, the font_family_without_font_face bug. Now both family and filename normalize away all non-alphanumerics; families match longest-key-first so a parent family can't swallow a more specific one's files (TT Norms Pro vs Mono); each file is claimed once; "demibold" ranks before "bold"; and when nothing matches it warns loudly at build time instead of returning "". Also: parseFonts() falls back to h1/h2/title/hero display roles, and the frame-worker + caption authoring docs spell out that only shipped font files render — no system CJK/Devanagari families on the headless renderer. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(hyperframes-media): enforce sign-in preflight on standalone BGM/TTS A one-off "generate me a BGM" request went straight to local MusicGen without recommending sign-in: bgm.md/tts.md framed the no-credential path as an automatic fallback, so the generation path bypassed the Preflight stop, and the preflight used a bare `hyperframes auth status` that isn't on PATH in a fresh `npx skills` project. - Preflight now applies to one-off generation as well as workflows, uses `npx hyperframes auth status`, and says: if the CLI can't run, still recommend signing in and STOP — never treat "no credential" as a silent green light for local generation. - bgm.md and tts.md point at the Preflight before generating, reframing local generation as the fallback the user opts into, not a default. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(auth): add Authentication & API keys guide Document signing in, the keys each capability (voice, music, capture) uses, their resolution priority, and the fully local fallback. Add the guide to the nav and cross-link it from the cloud deploy note and the CLI env-var reference. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(lint): strip HTML comments in a fixpoint loop (CodeQL) Single-pass <!-- --> removal can re-form a complete comment from adjacent markers (e.g. `<<!-- -->!-- ... -->`), letting a decoy <template> survive and hijack the template-boundary match. Loop to a fixpoint, mirroring the captions.mjs precedent; add a regression test that fails on single-pass (2 root findings) and passes on the loop. Also wrap the build-frame.mjs node:fs imports to satisfy oxfmt — the new copyFileSync import pushed the line past the width limit, which was the sole cause of the Format / Preflight CI failures. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(lint): strip HTML comments with a linear scan (CodeQL ReDoS) The fixpoint loop still ran a /<!--[\s\S]*?-->/ regex per pass, which backtracks O(n^2) on inputs with many unterminated "<!--" — CodeQL js/polynomial-redos (high). Looping the same regex (the prescribed fix) never addressed this; only the regex itself does. Replace it with an indexOf-based linear strip in utils.ts (stripHtmlComments), kept in a fixpoint loop so markers that re-form when a comment is removed are still stripped. 200k unterminated "<!--" now strips in ~3ms instead of quadratic time; behavior is otherwise unchanged — unterminated comments are kept verbatim, as the old regex left them. The re-forming regression test still guards it. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(auth): make TTS/BGM sign-in guidance accurate and runnable From team review of the not-signed-in onboarding: - OAuth is a `hyperframes auth login` feature only. The separate `heygen` CLI is API-key-only — `heygen auth login` stores a pasted key, it is not OAuth and does not create an account. Stop presenting the two CLIs as the same OAuth/sign-up step. - Use `npx hyperframes` in every imperative and runtime hint. Bare `hyperframes` is not on PATH on a fresh machine (command not found); only `npx hyperframes` is guaranteed. Also updates the JSON recommended_action. - Drop `heygen auth login` from the terminal/skill onboarding: it needs its own install and there is no `npx heygen`, so it was a command-not-found trap. The shared-credential fact stays in the reference docs. Covers the `auth status` guidance + tests, the Authentication docs, the shared hyperframes-media preflight (SKILL, requirements, tts, error hints), and the `npx hyperframes auth status` preflight in every TTS/BGM workflow (pr-to-video, product-launch-video, faceless-explainer, website-to-video, music-to-video). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.8
parent
c7b9bf3386
commit
54cab331d0
@@ -15,8 +15,9 @@
|
||||
// node captions.mjs build --storyboard ./STORYBOARD.md --audio-meta ./audio_meta.json --hyperframes . --out ./caption_groups.json
|
||||
//
|
||||
// CAPTION LOOK — two sources, picked automatically:
|
||||
// 1. PRESET SKIN (preferred). If a project-local `caption-skin.html` exists (Step 2
|
||||
// copies the chosen frame-preset's skin into the project), it is the caption look.
|
||||
// 1. PRESET SKIN (preferred). If a project-local `.hyperframes/caption-skin.html`
|
||||
// exists (Step 2 copies the chosen frame-preset's skin into the project), it is
|
||||
// the caption look.
|
||||
// It is a brand-token-strict skin with three reserved holes; this script fills them
|
||||
// and wraps the result in a <template> for the engine:
|
||||
// - `var GROUPS = [];` → the computed caption groups
|
||||
@@ -68,7 +69,12 @@ function runBuild(argv) {
|
||||
const outPath = resolve(flag(argv, "out", join(hyperframesDir, "caption_groups.json")));
|
||||
const htmlPath = join(hyperframesDir, "compositions/captions.html");
|
||||
const overridesPath = join(hyperframesDir, "caption-overrides.json");
|
||||
const skinPath = resolve(flag(argv, "skin", join(hyperframesDir, "caption-skin.html")));
|
||||
const skinArg = flag(argv, "skin", null);
|
||||
const hiddenSkinPath = join(hyperframesDir, ".hyperframes", "caption-skin.html");
|
||||
const legacySkinPath = join(hyperframesDir, "caption-skin.html");
|
||||
const skinPath = resolve(
|
||||
skinArg ?? (existsSync(hiddenSkinPath) ? hiddenSkinPath : legacySkinPath),
|
||||
);
|
||||
const framePath = resolve(flag(argv, "frame", join(hyperframesDir, "frame.md")));
|
||||
|
||||
if (!existsSync(storyboardPath)) die(`STORYBOARD.md not found at ${storyboardPath}`);
|
||||
@@ -291,8 +297,8 @@ function brandFontFaces(framePath, hyperframesDir) {
|
||||
const weightOf = (n) => {
|
||||
const s = n.toLowerCase();
|
||||
if (/black|heavy|ultra|extrabold/.test(s)) return 800;
|
||||
if (/semibold|demibold/.test(s)) return 600; // before /bold/ — "demibold" contains "bold"
|
||||
if (/bold/.test(s)) return 700;
|
||||
if (/semibold|demibold/.test(s)) return 600;
|
||||
if (/medium/.test(s)) return 500;
|
||||
if (/light|thin/.test(s)) return 300;
|
||||
return 400; // book / regular / roman
|
||||
@@ -305,10 +311,22 @@ function brandFontFaces(framePath, hyperframesDir) {
|
||||
: /\.ttf$/i.test(f)
|
||||
? "truetype"
|
||||
: "opentype";
|
||||
// Normalize away ALL non-alphanumerics (spaces, underscores, hyphens) on BOTH the
|
||||
// family name and the filename. Real font files use "_" / "-" as word separators
|
||||
// ("TT_Norms_Pro_Bold.woff2"), so stripping only whitespace never matched them — the
|
||||
// family key "ttnormspro" failed `startsWith` against "tt_norms_pro_bold", and the
|
||||
// function silently returned "" → captions shipped with NO @font-face for any
|
||||
// underscore/hyphen-named brand font (e.g. TT Norms Pro), which is exactly the
|
||||
// font_family_without_font_face bug.
|
||||
const norm = (s) => s.toLowerCase().replace(/[^a-z0-9]/g, "");
|
||||
const faces = [];
|
||||
const seen = new Set();
|
||||
for (const fam of families) {
|
||||
const key = fam.replace(/\s+/g, "").toLowerCase();
|
||||
const claimed = new Set(); // each file is claimed by the MOST SPECIFIC family only
|
||||
// Match the longest family key first so "TT Norms Pro" can't swallow the files that
|
||||
// belong to "TT Norms Pro Mono" (its key is a prefix of the longer one's).
|
||||
const ranked = [...families].sort((a, b) => norm(b).length - norm(a).length);
|
||||
for (const fam of ranked) {
|
||||
const key = norm(fam);
|
||||
for (const d of dirs) {
|
||||
let files = [];
|
||||
try {
|
||||
@@ -318,17 +336,34 @@ function brandFontFaces(framePath, hyperframesDir) {
|
||||
}
|
||||
for (const f of files.sort()) {
|
||||
if (!/\.(woff2|woff|ttf|otf)$/i.test(f)) continue;
|
||||
if (!f.replace(/\s+/g, "").toLowerCase().startsWith(key)) continue;
|
||||
if (claimed.has(f)) continue; // a more specific family already took this file
|
||||
if (!norm(f.replace(/\.(woff2|woff|ttf|otf)$/i, "")).startsWith(key)) continue;
|
||||
const w = weightOf(f);
|
||||
const dedup = `${fam}-${w}`;
|
||||
if (seen.has(dedup)) continue; // one src per weight; assets/fonts wins over capture
|
||||
seen.add(dedup);
|
||||
claimed.add(f);
|
||||
faces.push(
|
||||
` @font-face { font-family: '${fam}'; src: url('${d.rel}/${f}') format('${fmtOf(f)}'); font-weight: ${w}; font-display: block; }`,
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
// Loud signal instead of a silent "". If frame.md named a brand font but no file
|
||||
// matched, the caption text WILL fall back to a generic font in the render — surface
|
||||
// the cause here (at build time) rather than letting it surface 2 steps later as a
|
||||
// font_family_without_font_face lint error disconnected from its root cause.
|
||||
if (!faces.length) {
|
||||
const where = dirs.length
|
||||
? dirs.map((d) => d.rel).join(" / ")
|
||||
: "assets/fonts or capture/assets/fonts (neither exists)";
|
||||
console.warn(
|
||||
` ⚠ captions: frame.md names font ${families.map((f) => `"${f}"`).join(", ")} ` +
|
||||
`but no matching .woff2/.woff/.ttf/.otf was found in ${where} — captions will fall back ` +
|
||||
`(text may render in the wrong font). Stage a font file whose name starts with the family ` +
|
||||
`(e.g. "TT Norms Pro" → TT_Norms_Pro_Bold.woff2) so it ships with the project.`,
|
||||
);
|
||||
}
|
||||
return faces.join("\n");
|
||||
}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user