Files
hyperframes/skills/hyperframes-media/references/bgm.md
T
WaterrrForeverandClaude Opus 4.8 54cab331d0 feat(cli): shared TTS/BGM auth preflight + caption and skill-workflow fixes (#1697)
* fix: handle caption skin workflow

* docs(skills): simplify the finalize step across video workflows

- Drop --strict-layout; all skills use plain `hyperframes inspect`
- Add the caption text_box_overflow false-positive note to faceless-explainer
- On a failed check, the orchestrator makes the cheapest safe edit itself
  (no worker re-dispatch / Step 3 backtrack language)
- Snapshot: glance at the stitched contact-sheet.jpg and move on

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(auth): onboarding-first `auth status` + shared TTS/BGM preflight

When no HeyGen credential is configured, `hyperframes auth status` now
prints registration-first guidance instead of a terse error:

- Interactive / agent-driven sessions get sign-in guidance led by
  `hyperframes auth login` (the OAuth step that also creates an account
  and is shared with heygen-cli), and never steer users to a per-repo
  `.env`. CI / non-interactive runs get a terse note. Exit 1 is kept so
  the "am I logged in?" `$?` contract still holds.
- It probes which local engine voice/music will fall back to (Kokoro /
  MusicGen, mirroring the skill resolution order) and whether their
  Python deps are installed, with a pip hint when missing. `--json`
  exposes `recommended_action` + `offline_engines` for skills to branch.
- `doctor` gains matching "TTS (Kokoro)" / "BGM (MusicGen)" checks via
  the same shared probe (findPython/hasPythonModules extracted to
  tts/python.ts; provider resolution in audio/providers.ts).

Every TTS/BGM workflow now relays this at Step 0 (setup) instead of
improvising its own "missing key" prompt: pr-to-video, product-launch-
video, faceless-explainer, website-to-video, music-to-video. The
canonical behavior + key-priority table live once in hyperframes-media.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(pr-to-video): scale recommended video length to PR change size

Step 0 led with a fixed ~60-90s length default. Now the recommended
length is derived from the PR's diff stat (lines added+deleted, nudged
by file count) on a tier scale (trivial ~20-40s → large ~110-180s, hard
cap ~3 min), reusing the same PR peek already done to infer the angle.
The agent states the basis when proposing it, and a huge PR with one
headline change still stays tight. User can always override.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(captions): embed brand fonts whose files use separators

brandFontFaces() matched font files by stripping only whitespace, so an
underscore/hyphen-named file (TT_Norms_Pro_Bold.woff2) never matched the
family key "ttnormspro" — captions shipped with no @font-face, the
font_family_without_font_face bug. Now both family and filename normalize
away all non-alphanumerics; families match longest-key-first so a parent
family can't swallow a more specific one's files (TT Norms Pro vs Mono);
each file is claimed once; "demibold" ranks before "bold"; and when
nothing matches it warns loudly at build time instead of returning "".

Also: parseFonts() falls back to h1/h2/title/hero display roles, and the
frame-worker + caption authoring docs spell out that only shipped font
files render — no system CJK/Devanagari families on the headless renderer.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(hyperframes-media): enforce sign-in preflight on standalone BGM/TTS

A one-off "generate me a BGM" request went straight to local MusicGen
without recommending sign-in: bgm.md/tts.md framed the no-credential path
as an automatic fallback, so the generation path bypassed the Preflight
stop, and the preflight used a bare `hyperframes auth status` that isn't
on PATH in a fresh `npx skills` project.

- Preflight now applies to one-off generation as well as workflows, uses
  `npx hyperframes auth status`, and says: if the CLI can't run, still
  recommend signing in and STOP — never treat "no credential" as a silent
  green light for local generation.
- bgm.md and tts.md point at the Preflight before generating, reframing
  local generation as the fallback the user opts into, not a default.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(auth): add Authentication & API keys guide

Document signing in, the keys each capability (voice, music, capture)
uses, their resolution priority, and the fully local fallback. Add the
guide to the nav and cross-link it from the cloud deploy note and the
CLI env-var reference.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(lint): strip HTML comments in a fixpoint loop (CodeQL)

Single-pass <!-- --> removal can re-form a complete comment from
adjacent markers (e.g. `<<!-- -->!-- ... -->`), letting a decoy
<template> survive and hijack the template-boundary match. Loop to a
fixpoint, mirroring the captions.mjs precedent; add a regression test
that fails on single-pass (2 root findings) and passes on the loop.

Also wrap the build-frame.mjs node:fs imports to satisfy oxfmt — the
new copyFileSync import pushed the line past the width limit, which
was the sole cause of the Format / Preflight CI failures.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(lint): strip HTML comments with a linear scan (CodeQL ReDoS)

The fixpoint loop still ran a /<!--[\s\S]*?-->/ regex per pass, which
backtracks O(n^2) on inputs with many unterminated "<!--" — CodeQL
js/polynomial-redos (high). Looping the same regex (the prescribed
fix) never addressed this; only the regex itself does.

Replace it with an indexOf-based linear strip in utils.ts
(stripHtmlComments), kept in a fixpoint loop so markers that re-form
when a comment is removed are still stripped. 200k unterminated
"<!--" now strips in ~3ms instead of quadratic time; behavior is
otherwise unchanged — unterminated comments are kept verbatim, as the
old regex left them. The re-forming regression test still guards it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(auth): make TTS/BGM sign-in guidance accurate and runnable

From team review of the not-signed-in onboarding:

- OAuth is a `hyperframes auth login` feature only. The separate `heygen`
  CLI is API-key-only — `heygen auth login` stores a pasted key, it is not
  OAuth and does not create an account. Stop presenting the two CLIs as the
  same OAuth/sign-up step.
- Use `npx hyperframes` in every imperative and runtime hint. Bare
  `hyperframes` is not on PATH on a fresh machine (command not found); only
  `npx hyperframes` is guaranteed. Also updates the JSON recommended_action.
- Drop `heygen auth login` from the terminal/skill onboarding: it needs its
  own install and there is no `npx heygen`, so it was a command-not-found
  trap. The shared-credential fact stays in the reference docs.

Covers the `auth status` guidance + tests, the Authentication docs, the
shared hyperframes-media preflight (SKILL, requirements, tts, error hints),
and the `npx hyperframes auth status` preflight in every TTS/BGM workflow
(pr-to-video, product-launch-video, faceless-explainer, website-to-video,
music-to-video).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 23:20:39 +08:00

7.1 KiB
Raw Blame History

Background music (BGM)

One music bed per composition, produced by the shared audio engine (scripts/audio.mjsscripts/lib/bgm.mjs). Two routes, chosen by the engine's one switch — whether a HeyGen credential is present:

  • HeyGen retrieval — the default when credentialed. Search HeyGen's music catalog by mood, download the top track. No generation; same ~/.heygen / $HEYGEN_API_KEY credential as TTS.
  • Local generation (Lyria → MusicGen) — the fallback when there is no credential (or when asked for explicitly). Generate a WAV from a mood prompt. There is no npx hyperframes bgm command; the engine spawns scripts/lyria-recipe.py or an inline MusicGen script directly.

Run the Preflight first — no credential is not a green light to silently generate locally. Before generating, complete the sign-in Preflight (see ../SKILL.md → Preflight): run npx hyperframes auth status, recommend signing in, and STOP for the user's choice (sign in for HeyGen's music library, or continue offline with local generation). This applies to a one-off "generate a BGM" request just as much as inside a full workflow.

Driving it from the request

audio_request.jsonbgm: { mode?, query?, prompt? }:

  • moderetrieve | generate | none. Omit for auto (retrieve when credentialed, else generate). An explicit retrieve is strict: no credential ⇒ skip, never a detached generate (so a caller with no wait-bgm step, e.g. product-launch, can't get a pending job it won't await).
  • query — the mood, used for retrieval and as a fallback prompt seed (e.g. a storyboard's music: field, falling back to messagearc"calm cinematic underscore").
  • prompt — an explicit full prompt for generation; omit and the engine infers one (see Mood inference). Optional blob / archetype / arc feed that inference.

HeyGen retrieval (default)

searchSounds(query, "music", { limit: 5 })GET /audio/sounds?query=<mood>&type=music&limit=5. Take the top result (ranked by score), download its presigned audio_urlassets/bgm/track.mp3. Synchronous. No match → skip (BGM is optional; never fail the render over it). Cue written to audio_meta.json:

{
  "path": "assets/bgm/track.mp3",
  "volume": 0.8,
  "mode": "retrieve",
  "query": "calm cinematic underscore",
  "duration_s": 42.0,
}

volume is 0.8 under narration, 0.9 for a silent film (no voice). bgm_pending is false — the file is on disk when the engine returns.

Local generation (fallback) — Lyria → MusicGen

Spawned detached so voice work isn't blocked; audio_meta.bgm_pending: true and bgm_pid / bgm_log are set until it finishes. Run scripts/wait-bgm.mjs before assembling — it polls the output file / process / log, detects crashes, and writes bgm_status.json (status: ready | failed | timeout | disabled). A failed/absent track is simply omitted; it never blocks voice/SFX.

Order Provider Env / deps Speed Quality
1 Google Lyria RealTime $GEMINI_API_KEY or $GOOGLE_API_KEY + google-genai (auto-installed on demand) Real-time stream (≈ requested duration) Production-grade
2 MusicGen (facebook/musicgen-small) Python transformers + torch + soundfile + numpy (~300 MB first run; auto-installed) Slow on CPU; fast on Apple MPS / CUDA Decent; prompt-only control

Output → assets/bgm/track.wav, target = total voice duration. MusicGen generates one seed clip (≤2830s, under the decoder's positional limit) then crossfade-loops it up to the target (or trims down if shorter), avoiding per-segment seams. Backend selection is by what can actually run: Lyria only when import google.genai succeeds, else MusicGen; if neither can be made to run, BGM is skipped (voice + SFX still render).

Mood inference (the generate prompt)

inferBgmPrompt() in scripts/lib/bgm.mjs: an explicit prompt wins; otherwise industry-keyword base → narrative-archetype shape → emotional-arc tiebreaker.

Match in blob / query Base prompt BPM
crypto / nft / web3 / defi / token / blockchain atmospheric electronic, deep bass, futuristic synths, restrained percussion 100
finance / fintech / bank / payment / invest / wealth calm cinematic, soft strings, subtle piano, restrained percussion 92
creative / agency / design / studio / art / brand playful electronic, warm pads, light percussion 115
(default: SaaS / tech / platform) uplifting corporate tech, bright modern piano with synth pads 108

Archetype then reshapes the arc — PAS → "MINOR to MAJOR" build; BAB / future-pacing → aspirational rising; feature-cascade → +10 BPM driving; demo-loop → 8 BPM minimal. The emotional arc breaks remaining ties (tension→relief, excitement, trust/reassurance).

Lyria knobs (direct recipe use)

The engine bakes BPM / scale into the prompt text (via the inference above) and passes only --output / --duration / --prompt to the recipe. If you invoke scripts/lyria-recipe.py directly you can also set: --bpm (90110 calm, 110130 energetic), --brightness (01, ≥0.7 promotional), --density (01, higher = fuller), --scale (MAJOR / MINOR / PENTATONIC / …), --negative-prompt (styles to exclude). MusicGen ignores all of these — put the mood in the prompt.

Failure modes

Failure Behavior
No music match (retrieve) bgm: null, anomaly logged. Render proceeds without BGM.
Explicit retrieve, no credential Skipped (no silent generate fallback). Use mode: generate or omit mode for auto.
Neither Lyria nor MusicGen can run (generate) bgm disabled with a pip install … hint. Voice + SFX still render.
Generate still rendering at assemble time bgm_pending: true; wait-bgm.mjs waits/checks and writes bgm_status.json first.
Generate crashed wait-bgm.mjsbgm_status.json { status: "failed" }; the <audio> track is omitted.

BGM failure never blocks a render.