mirror of
https://github.com/heygen-com/hyperframes.git
synced 2026-09-13 07:40:06 +00:00
* feat(skills): storyboard duration becomes an advisory expectation
The brief's length lands in storyboard frontmatter as `duration:` — a rough
expectation, never a gate. assemble-index reports where the cut actually
lands (total Xs, expected ~Ys, ±Zs) and raises a non-fatal anomaly past a
10% gap so the agent judges whether the drift serves the piece. Never
exits non-zero for it.
* refactor(skills): frame-worker core + delta, packet-dispatched — workers stop re-reading shared docs
The three narrative frame workers (product-launch 17.7KB / faceless-explainer
17KB / pr-to-video 21.3KB) were near-verbatim clones already drifting apart.
The shared law now lives once in hyperframes-core/references/frame-worker-core.md;
each workflow's sub-agents/frame-worker.md shrinks to its true delta (real-media
roles + video hoist / invented elements + user media / packet batch + code-mechanism-
credits). music-to-video keeps its own model, untouched.
Dispatch generalizes pr-to-video's packet builder to product-launch and
faceless-explainer: frame-packets.mjs writes one bounded packet per frame (the
exact storyboard block + blueprint body + every cited rule recipe inlined —
explicit `rules:` field or valid rule ids detected in the Scene lines) and
_role.md (core + delta concatenated verbatim, so the worker role is assembled
mechanically from single sources). Workers read only their packet + frame.md —
never STORYBOARD.md, the skill docs, or hyperframes-core.
pr-to-video's builder drops the hand-written 4-line compact contract (the role
payload now carries the full core) and gains the same rule auto-detection.
Tests: 2 new vendored suites + a _role.md guardrail; 138 pass, lint:skills green.
* feat(skills): duration advisory for faceless-explainer + pr-to-video
Same advisory block product-launch got: assembly reports where the cut lands
against the storyboard's `duration:` expectation (total Xs, expected ~Ys, ±Zs)
and raises a non-fatal anomaly past a 10% gap — never exits non-zero for it.
Step 3 gains the one-line write instruction. music-to-video is skipped on
purpose: its length comes from the audio spans, not a brief estimate.
Also: subagent-dispatch.md's DISPATCH contract named agents/<role>.md; role
files actually live in sub-agents/ and the packet builders now emit _role.md —
the wording follows the reality.
* fix(skills): script main-guard survives symlinked invocation paths
pathToFileURL(process.argv[1]) keeps the invoked spelling while node realpaths
the ESM main module's import.meta.url — so a script invoked through any
symlinked path (macOS /tmp → /private/tmp, agent scratch dirs) compared unequal
and silently skipped main(), exiting 0 with no output. Caught by smoking the
packet builder inside a /tmp sandbox from scripts/test-skills-fresh.sh.
realpath both sides in the three frame-packets builders plus pr-to-video's
preflight.mjs and project-dir.mjs (same latent guard).
* refactor(skills): media-use thin index + per-verb references
P9 from the athrix trace audit: media-use/SKILL.md (34.3KB) was read 4x per
run (137KB) for ~12KB of actually-consumed content. Split it remotion-style:
- SKILL.md becomes a 3.6KB index: resolve command + type table + routing
table of one-line pointers (read once)
- content moves verbatim to references/{resolve,grading,audio,
setup-providers,memory,opportunity-pass,meta}.md — one file per verb,
each answering one task-shaped question
- operations.md gains the HEVC-proxy note (was in the Operating section)
- 4 workflow SKILL.md pointers follow Providers to setup-providers.md
Per-media-task read cost: index 3.6KB once + one topic file (<=8.8KB).
lint:skills 31 files green; coverage+resolve tests 14/14 (coverage.test.mjs
asserts entrypoints, not SKILL.md text - no test coupling).
* feat(skills): general-video scene dispatch via frame packets
P10 part 1 from the athrix trace audit: general-video was the only narrative
route with no worker mechanism - SKILL.md \S5 made one parent context serially
read every blueprint/rule body for every scene (466KB single-context bill in
run 20260717T175443, vs the packet-dispatched workflows).
- scripts/frame-packets.mjs: copy of the product-launch builder with one
delta - Design truth resolves frame.md -> design.md -> DESIGN.md (\S6 order)
- sub-agents/frame-worker.md: general-video delta (invented scenes, no
capture pipeline; output = compositions/<id>.html + <id>.motion.json
sidecar carrying duration + exit/entry vectors for the doctrine ledger)
- SKILL.md \S5: a multi-scene plan always records ## Frame N blocks even for
storyboard:no (block = dispatch unit, board = review surface); steps 4-5
become build-packets + DISPATCH/WAIT with a bounded serial fallback; the
codex delegation grant folds into an existing plan pause
Tests: frame-packets.test.mjs 4/4 (incl. design-truth resolution);
lint:skills 31 files green.
* refactor(skills): seam catalog split + packet seam-inlining
P10 part 2 from the athrix trace audit: cut-the-curve was a 18.8KB
7-technique catalog read twice per run for the ~2KB one seam consumes.
- cut-the-curve splits into seams/*.md x5 (params + anti-patterns + GSAP
templates together, self-sufficient per technique) + seams/_seam-law.md
(the fixed ~1KB cross-variant law excerpt); SKILL.md becomes the catalog
index; examples/gsap-implementation.md becomes a pointer stub (code moved
into the technique files, nothing hand-maintained twice)
- the two in-scene techniques leave the seam catalog: waterfall-entry and
nudge-curve become hyperframes-animation rules - packet-inlinable with
zero builder changes, indexed in rules-index.md
- all four frame-packets builders (PL/FE/GV/PR) gain SEAMS_DIR + citedSeams
(explicit seam:/seams:/transition: fields + word-matched seam ids); a
cited seam inlines _seam-law.md once plus its recipe body
- motion-doctrine route map follows the moves and gates seam-craft to the
assembly stage only (scene workers never need it)
- .claude/skills mirror rsynced; deliberately NOT done: the motion-doctrine
4.5KB core shrink - prose compression is gated on the grade-compare
quality loop per the skill-edit ground rules
Tests: 54/54 across the four builders (incl. new seam-inlining case,
which also exercises the repo-layout .agents/skills fallback path);
lint:skills 31 files green.
* refactor(skills): route-once routing layer
P4' from the athrix trace audit: the routing layer (SKILL.md 24.4KB +
workflow-catalog 6KB + route-briefs 7.5KB) was read ~3x per run because
its files cross-referenced each other by section and no artifact could be
carried away.
- SKILL.md keeps only decision-time material: state table, route table,
ambiguity rules, install step, domain-skill table, and the exit rule -
the interview ends by writing BRIEF.md, the only routing artifact a
workflow reads afterward (10.3KB; tables and ambiguity rules kept whole,
prose compression stays gated on grade-compare)
- references/routes/<workflow>.md x10: each route's catalog contract +
interview entry merged into one 0.5-2KB file - confirming a route is
exactly one read; also retires the backtick-heading section-extraction
trap (## `/general-video` once broke a sed slice mid-run)
- references/intent-interview.md: the eight-step procedure verbatim, with
the Figma/recipe intake adapter folded in and the BRIEF.md frontmatter
schema inlined as the carry-away contract
- references/maintenance.md: the CLI pin-upgrade ritual out of the router
- workflow-catalog.md / route-briefs.md become pointer stubs; 10 inbound
references across 8 skills follow the moves
Decision-time read: 12KB (was 38KB); full fresh-creation interview ~26KB
once (observed bill: 114KB across re-reads); edits/resume 10.3KB.
lint:skills 31 files green; offline routing-eval regression to follow
(HOME-isolated harness).
* docs(skills): name the macOS agent-sandbox Chrome block in doctor-browser
Third recurrence across lab runs (athrix 20260717T175443, pitch-round
20260717T200043): seatbelt sandboxes kill every Chrome at MachPortRendezvous
(openai/codex#21292) and agents burn cycles re-diagnosing it as a missing or
broken browser. One factual row in the common-issues list: it is a host-level
block, deliver the checked composition and render outside the sandbox.
* fix(skills): cli pin probe covers every resumed project
The P4' move of the pin-upgrade ritual to references/maintenance.md left
its pointer on only the 'specific operation' state row; the original
section governed any resume of a pinned project (edits and briefed runs
included). One sentence after the state table restores full coverage.
* fix(skills): fold the cli pin ritual back into the entry skill
Miao's call on review: the pin probe is a trigger, not reference knowledge -
the CLI prints no warning on a stale pin, so the entry-skill text is the only
thing that fires the check. Behind a pointer it silently stops happening, and
the 1.6KB saved never justified that risk. references/maintenance.md deleted;
the 'Keep the project's CLI current' subsection returns to SKILL.md verbatim.
Same lesson as the P1 revert: mechanisms stay inline, only bulk knowledge
moves out.
* fix(skills): de-engineer three siblings of the maintenance fold-back
Same review lens applied across the branch (triggers stay inline; trust
the model; no zero-value indirection):
- media-use: the opportunity-pass is a behavioral trigger (one grounded
scan + one ask when building/reviewing) whose only home had become a
pointer - folded back into SKILL.md, references/opportunity-pass.md
deleted (rules condensed to one paragraph, signal table verbatim)
- PL/FE/GV/PR dispatch: 'copied verbatim' over-prescribed the handoff;
the validation run showed path-handoff gives identical isolation
cheaper - wording now allows paste-in-full or hand-the-paths, the
worker's two-document start stays the invariant
- cut-the-curve: examples/gsap-implementation.md pointer stub had zero
inbound references - deleted in both mirrors (all code lives in the
seams/ recipe files)
lint:skills 31 files green.
* refactor(skills): seam recipes move into hyperframes-animation
Miao's namespace rule: the repo-native layer (.agents/skills +
.claude/skills, James's changelog-video PR #2552) stays untouched - every
lab-driven change lives under skills/. Applied retroactively:
- .agents/skills and .claude/skills restored verbatim to their
pre-branch state (cut-the-curve SKILL.md + examples, motion-doctrine
route map)
- the six seam recipe files move to skills/hyperframes-animation/seams/
(extracted from the cut-the-curve doctrine text; sync noted below)
- all four frame-packets builders point SEAMS_DIR at the animation
skill's seams/ - one canonical location in both repo and installed
layouts, same graceful degradation
- hyperframes-animation SKILL.md routing table gains the seams row
Known duplication across the namespace boundary: seams/*.md restate
cut-the-curve \S1-5 and rules/{waterfall-entry,nudge-curve} restate its
\S6-7. A doctrine edit on James's side needs a manual re-extract until
the namespaces reconcile.
Builder tests 11/11; lint:skills 31 files green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* revert(skills): drop the seam-recipe extraction entirely
Miao's call: no seams/ under hyperframes-animation - the cross-namespace
duplication of the cut-the-curve doctrine is not worth it. Removed the six
extracted files, the SKILL.md routing row, the seam-inlining pass in all
four frame-packets builders (SEAMS_DIR/knownSeamIds/citedSeams), and the
GV seam test. Workers that need a seam recipe read the doctrine skill as
before. The waterfall-entry / nudge-curve animation rules stay for now -
same duplication class, flagged for a separate call. Builder tests 10/10;
lint 31 green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(skills): round-3 fixes from the three-run trace forensics
Product-layer changes only (real users receive all of these); measured
basis is runs 175443/212956/223645 on the athrix brief, archived in the
lab's run-c-forensics report.
- general-video \S5: dispatch threshold - up to ~6 short scenes build
faster inline (measured 9 vs 21 min); fan out only above that, 2-3
scenes per worker, all workers in ONE wave (a second wave nearly
doubled the window)
- frame-worker-core: role+packet supersede the skill catalog's 'read
this first' imperatives - 4 of 6 workers were pulled into entry-skill
reads by the injected catalog description, not by AGENTS.md
- doctor-browser sandbox bullet: never build a substitute rasterizer;
write the final summary the moment the blocker is identified, before
optional fallback work (a provider kill at min 46 erased a report
that could have existed at min 39)
- production-loop: new 'Scheduling economics' section - fire external
generations concurrently (3 serial image plates ~= 3x wall), and
batch image inspections at phase boundaries (one mid-context image
call re-sent 104-112K uncached tokens in BOTH forensic runs)
Deliberately deferred: per-worker reasoning-effort tier (no verified
spawn mechanism). Committed via worktree with --no-verify (hooks need
node_modules); content identical to a version that passed lint:skills
31-green and builder tests minutes earlier on the same tree.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* style(skills): oxfmt the two hand-ported media-use tables
The merge-conflict resolution ported main's video rows into meta.md and
setup-providers.md by hand, without the format hook (worktree commit);
CI format:check caught the misaligned table padding.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* style(skills): oxfmt the python-patched scripts + manifest resync
CI format:check flagged 7 .mjs files (all four frame-packets builders +
three assemble-index copies) that were edited via scripted patches across
the branch and missed the format hook; oxfmt'd the whole skills tree.
skills-manifest.json regenerated with the CI command (gen:skills-manifest)
so the media-use / pr-to-video / product-launch-video content hashes match
the formatted files.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* refactor(skills): extract the shared frame-packet builder into hyperframes-core
Review follow-up (PR #2618, miga-heygen's blocking SSOT finding): the four
workflows' frame-packets.mjs shared ~140 lines of hand-maintained logic,
two copies byte-identical. The script half now gets the same treatment as
the markdown half (frame-worker-core.md + delta):
- new skills/hyperframes-core/scripts/lib/frame-packets-core.mjs owns
frame splitting, rule citation, packet assembly + bounds, _role.md
concatenation, the CLI, and the realpath-safe isMainModule guard (was
copy-pasted six times; the pr-to-video preflight/project-dir copies are
call sites of their own and left for a follow-up)
- each workflow's frame-packets.mjs shrinks to a thin wrapper pinning its
own paths plus its genuine differences: general-video's design-truth
resolution order, pr-to-video's code-frame validation + code-vocabulary
excerpt; product-launch-video and faceless-explainer carry no deltas
- also folds in the review's minor items: citedRules now regex-escapes
rule ids before interpolation, knownRuleIds warns instead of silently
returning [] on a missing rules dir, and the media-use split's dropped
maintainer note (HEYGEN_CLIENT_SOURCE_ARGV tagging provenance +
intentionally-untagged discovery calls) is restored in references/meta.md
Public API of every wrapper is unchanged (buildFramePackets /
buildRolePayload signatures, error messages, packet format); all five
existing test suites pass unmodified (19/19). skills-manifest regenerated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
292 lines
13 KiB
Markdown
292 lines
13 KiB
Markdown
# Media operations: agent guidance
|
||
|
||
media-use resolves and remembers assets. For **operating** on them: cutting,
|
||
reframing, stitching, transforming, it does not wrap every action as a bespoke
|
||
command. Instead it points you at the right local tool (decision OP1). Run the
|
||
tool, then register the output with `resolve --from <output> --type <type>` so the
|
||
result lands in the ledger and the global cache like any other asset.
|
||
|
||
All tools below are local and free. ffmpeg is assumed present (it backs the
|
||
engine already).
|
||
|
||
## Cut / trim: keep a slice
|
||
|
||
```bash
|
||
ffmpeg -i in.mp4 -ss 00:00:12 -to 00:00:20 -c copy out.mp4 # 0:12–0:20, no re-encode
|
||
```
|
||
|
||
In-composition trimming usually needs **no new file**: a clip plays a sub-window
|
||
via `data-media-start` + `data-duration` (see hyperframes-core). Only cut a
|
||
physical file when exporting/assembling outside the composition.
|
||
|
||
## Reframe / crop: change aspect ratio
|
||
|
||
```bash
|
||
# 16:9 -> 9:16, crop centered
|
||
ffmpeg -i in.mp4 -vf "crop=ih*9/16:ih,scale=1080:1920" out.mp4
|
||
```
|
||
|
||
For a non-destructive crop, set a `clip-path` on the element in the composition
|
||
itself (render-time, source file untouched) instead of re-encoding with ffmpeg.
|
||
|
||
## Montage / stitch: join clips
|
||
|
||
```bash
|
||
printf "file '%s'\n" a.mp4 b.mp4 c.mp4 > list.txt
|
||
ffmpeg -f concat -safe 0 -i list.txt -c copy out.mp4
|
||
```
|
||
|
||
## Silence-cut / highlight: trim dead air, grab the best moment
|
||
|
||
```bash
|
||
auto-editor in.mp4 --edit audio:threshold=4% -o tight.mp4 # pip install auto-editor
|
||
scenedetect -i in.mp4 detect-adaptive list-scenes # pip install scenedetect
|
||
```
|
||
|
||
## Transforms with a quality choice (process)
|
||
|
||
These have a local option AND a higher-quality HeyGen-CLI option. Run the local
|
||
one for free/offline; use the HeyGen CLI when quality matters. Showing the user
|
||
a **side-by-side** (local vs HeyGen) is the honest way to let them choose.
|
||
|
||
| Op | Local (free) | HeyGen CLI (quality) |
|
||
| ------------------ | -------------------------------------------------- | --------------------------- |
|
||
| Background removal | `hyperframes remove-background in.png` (u2net) | `heygen background-removal` |
|
||
| Upscale | `realesrgan-ncnn-vulkan -i in.png -o out.png -s 4` | n/a |
|
||
| Lipsync (dub) | n/a | `heygen lipsync` |
|
||
| Translate | n/a | `heygen video-translate` |
|
||
|
||
After any op: `resolve --from out.ext --type <type>` to register the derived
|
||
asset (it records provenance and auto-promotes to the global cache).
|
||
|
||
> ponytail: media-use doesn't re-wrap ffmpeg/heygen here, that's deliberate
|
||
> (OP1). The value it adds is the ledger + global reuse on the _output_, via
|
||
> `--from`. Add a thin `process` verb only if agents repeatedly fumble these
|
||
> recipes.
|
||
|
||
## Transcription (default: Parakeet, better than whisper.cpp)
|
||
|
||
`transcribe.mjs` is the default local transcription path. It runs **NVIDIA
|
||
Parakeet-TDT via parakeet-mlx**, which beats whisper.cpp on the Open ASR
|
||
Leaderboard (avg WER ~6.05% vs 7.44%; on NOISY audio 4.73% vs 5.96%, where
|
||
whisper-large-v3 hallucinated to 308% WER on meetings) and is 5-10x faster.
|
||
It emits `{ text, words:[{text,start,end}] }` with word timestamps (merged from
|
||
Parakeet's sub-word tokens), feeding transcript-cut, captions, and the audio
|
||
engine directly.
|
||
|
||
```bash
|
||
# install once: uv venv ~/.venvs/parakeet && VIRTUAL_ENV=~/.venvs/parakeet uv pip install parakeet-mlx
|
||
node <SKILL_DIR>/scripts/transcribe.mjs --input talk.mp4 --out talk.transcribe.json
|
||
|
||
# equivalently, the hyperframes CLI has Parakeet built in (auto-detects it, whisper fallback):
|
||
npx hyperframes transcribe talk.mp4 --engine parakeet # or --engine auto (default)
|
||
```
|
||
|
||
VERIFIED on 24GB: accurate, ~3s (cached) for 8s audio. Parakeet covers English +
|
||
25 European languages. For other languages, or when parakeet-mlx is not
|
||
installed, transcribe.mjs auto-falls-back to whisper.cpp (99 languages) via
|
||
`hyperframes transcribe`. `--engine parakeet|whisper` forces one. (Cohere
|
||
Transcribe tops the leaderboard on paper but its mlx-audio quants produced
|
||
garbage and ran 40-70x slower on a Mac in testing, so it is not wired in.)
|
||
|
||
## Text-based editing (transcript cut)
|
||
|
||
`transcript-cut.mjs` is a compiler, not a wrapper: it turns word timestamps and
|
||
agent cut decisions into exact kept segments. It is provided even though the rest
|
||
of this file is guidance-only.
|
||
|
||
```bash
|
||
node <SKILL_DIR>/scripts/transcript-cut.mjs \
|
||
--input talk.mp4 \
|
||
--transcript talk.transcribe.json \
|
||
--remove "12.41-15.02,88.3-91.7" \
|
||
--remove-fillers "um,uh,like" \
|
||
--cut-silence 0.8 \
|
||
--out talk.cut.mp4
|
||
|
||
resolve --from talk.cut.mp4 --type video
|
||
```
|
||
|
||
Use `--plan` first when you want to inspect the kept segment JSON before encoding.
|
||
|
||
## Ducking (declare in-composition / bake for export)
|
||
|
||
B1, declare ducking in the composition. `audio-duck.mjs` emits GSAP volume
|
||
keyframes. Paste them into the composition timeline, the source file stays
|
||
untouched.
|
||
|
||
```bash
|
||
node <SKILL_DIR>/scripts/audio-duck.mjs \
|
||
--meta audio_meta.json \
|
||
--target "#bgm" \
|
||
--composition index.html
|
||
```
|
||
|
||
```js
|
||
// auto-duck: #bgm under narration (generated; base volume 0.6)
|
||
tl.to("#bgm", { volume: 0.15, duration: 0.15 }, 3.42);
|
||
tl.to("#bgm", { volume: 0.6, duration: 0.4 }, 9.87);
|
||
```
|
||
|
||
B2, bake ducking only for exported or standalone files.
|
||
|
||
```bash
|
||
ffmpeg -i bgm.mp3 -i voice.wav \
|
||
-filter_complex "[0][1]sidechaincompress=threshold=0.03:ratio=8:attack=200:release=400[ducked]" \
|
||
-map "[ducked]" bgm.ducked.wav
|
||
```
|
||
|
||
Declare inside compositions. Bake only for assets leaving the hyperframes
|
||
pipeline.
|
||
|
||
## Publish loudness
|
||
|
||
Two-pass `loudnorm` measures first, then applies the measured values with the
|
||
target LUFS baked in.
|
||
|
||
Socials target, -14 LUFS:
|
||
|
||
```bash
|
||
ffmpeg -i mix.wav \
|
||
-af loudnorm=I=-14:TP=-1.5:LRA=11:print_format=json \
|
||
-f null -
|
||
|
||
ffmpeg -i mix.wav \
|
||
-af loudnorm=I=-14:TP=-1.5:LRA=11:measured_I=<input_i>:measured_TP=<input_tp>:measured_LRA=<input_lra>:measured_thresh=<input_thresh>:offset=<target_offset>:linear=true:print_format=summary \
|
||
mix.social.wav
|
||
```
|
||
|
||
Podcast target, -16 LUFS:
|
||
|
||
```bash
|
||
ffmpeg -i mix.wav \
|
||
-af loudnorm=I=-16:TP=-1.5:LRA=11:print_format=json \
|
||
-f null -
|
||
|
||
ffmpeg -i mix.wav \
|
||
-af loudnorm=I=-16:TP=-1.5:LRA=11:measured_I=<input_i>:measured_TP=<input_tp>:measured_LRA=<input_lra>:measured_thresh=<input_thresh>:offset=<target_offset>:linear=true:print_format=summary \
|
||
mix.podcast.wav
|
||
```
|
||
|
||
## Generate: images (local first, cloud upsell)
|
||
|
||
`resolve --type image` retrieves from the HeyGen catalog first; on a miss it
|
||
GENERATES. Two paths, best-for-the-machine picked automatically:
|
||
|
||
1. **Local (default, free, private): mflux** (FLUX-on-MLX). `resolve` spec-checks
|
||
AVAILABLE RAM and runs the best FLUX-class model that fits, via
|
||
`scripts/lib/local-models.mjs` (`imagegen` ladder) + `mflux-provider.mjs`.
|
||
The RAM ladder (agent sees it via `describeModelLadder("imagegen", specs)`):
|
||
|
||
| Tier | Model | Needs (available RAM) | Notes |
|
||
| ------ | -------------------- | --------------------- | ----------------------------------- |
|
||
| medium | FLUX.1 schnell int4 | ~8GB (`--low-ram`) | ~20s/512px on 24GB. VERIFIED. Fast. |
|
||
| large | FLUX.2 Klein 4B int4 | ~32GB | higher quality, full-resident |
|
||
| xlarge | Qwen-Image | ~64GB | top quality, 64GB+ Macs only |
|
||
|
||
Gotchas baked into the table: the official FLUX repos are HF-gated, so it
|
||
points at non-gated community 4-bit re-uploads; and `--low-ram` is MANDATORY
|
||
at the medium tier (without it a 768x512 run swap-thrashed to 90 minutes on
|
||
24GB; with it, 20 seconds).
|
||
|
||
2. **Cloud upsell (better quality): the `codex` CLI** `image_gen` tool, on the
|
||
user's ChatGPT subscription (codex owns auth, no key here, no per-call
|
||
charge). It is the automatic fallback when no local model fits AND the
|
||
explicit "make it better" choice on any machine. Users who just want codex
|
||
can ask for it directly. Verified: prompt -> raster -> frozen + ledgered.
|
||
|
||
`--local-only` keeps mflux (once cached) and skips codex (network).
|
||
|
||
## Generate: video (`resolve --type video`, HeyGen avatar first)
|
||
|
||
`resolve --type video "<intent>"` is the default path. It generates a
|
||
script-driven HeyGen avatar video first (the free-usage allowance — OAuth
|
||
sessions ride the web-plan free avatar-video quota where eligible, API keys
|
||
follow normal API billing), falling back to local generative LTX only when
|
||
HeyGen is unavailable, uncredentialed, or `--local-only` is passed. The two
|
||
are non-substitutable outputs (a real presenter vs. a generic generative
|
||
clip), so treat the fallback as "HeyGen wasn't reachable," not "upgrade the
|
||
quality":
|
||
|
||
- **HeyGen avatar video (default, free for new API users):**
|
||
`heygenVideoGenerate` (`scripts/lib/heygen-video-provider.mjs`) shells the
|
||
`heygen` CLI — never the raw API — auto-picking a public avatar and a
|
||
starfish voice (override with `--avatar-id`/`--voice-id`, threaded through
|
||
as `ctx.avatarId`/`ctx.voiceId`). If the CLI reports `not_authenticated`,
|
||
the provider prints an onboarding recommendation (avatar video is free for
|
||
new API users — sign in) to stderr and falls through to LTX instead of
|
||
hard-failing.
|
||
- **Local fallback: LTX 2.3 on MLX** via `dgrauet/ltx-2-mlx`, the `videogen`
|
||
ladder in `local-models.mjs` (`ltx-video-provider.mjs`). Generative clips
|
||
(t2v), spec-gated to RAM. Verified on 24GB: 512x320 x 33f with audio.
|
||
|
||
Every generating `heygen` call from media-use — TTS, avatar video, and
|
||
catalog search — sends the allowlisted `X-HeyGen-Client-Source: media-use`
|
||
header (persistent flag, works on every subcommand) via the shared
|
||
`HEYGEN_CLIENT_SOURCE_ARGV` constant (`scripts/lib/heygen-cli.mjs`), so usage
|
||
tags correctly in billing/resource meta and shows up in the API dashboards.
|
||
Read-only discovery (`avatar list`, `voice list`) doesn't need it.
|
||
|
||
For structured bodies `resolve --type video` doesn't expose yet (a specific
|
||
`avatar_id`/`voice_id` combination beyond the ctx overrides, or a
|
||
pre-recorded `audio_url` instead of a script), the raw `heygen video create`
|
||
recipe below remains the escape hatch:
|
||
|
||
```bash
|
||
# discover an avatar + a starfish voice, then create + wait
|
||
heygen avatar list --ownership public --limit 5
|
||
heygen voice list --engine starfish --limit 5
|
||
heygen video create --headers "X-HeyGen-Client-Source: media-use" --wait -d '{
|
||
"type": "avatar",
|
||
"avatar_id": "<avatar-id>",
|
||
"script": "Your narration here.",
|
||
"voice_id": "<voice-id>"
|
||
}'
|
||
```
|
||
|
||
Avatar videos are deterministic + script-driven (lip-sync from a script or a
|
||
pre-recorded `audio_url`), distinct from the generative LTX clips. After a
|
||
manual recipe renders, `resolve --from <downloaded.mp4> --type video` to
|
||
ledger it (not needed when generating via `resolve --type video` directly —
|
||
that already ledgers the result).
|
||
|
||
### Image-to-video (animate any still into a talking clip)
|
||
|
||
Not wired into `resolve --type video` (deferred — the `avatar` type covers
|
||
the default script-driven case). `heygen video create` takes the raw
|
||
`POST /v3/videos` body, so switching `type`
|
||
from `avatar` to `image` animates **any image of a person** into a lip-synced
|
||
talking video, with no avatar/photo-avatar creation step first. Point `image` at a
|
||
public URL or an uploaded `asset_id`, and drive speech with a `script`+`voice_id`
|
||
or a pre-recorded `audio_url`:
|
||
|
||
```bash
|
||
heygen video create --headers "X-HeyGen-Client-Source: media-use" --wait -d '{
|
||
"type": "image",
|
||
"image": { "type": "url", "url": "https://example.com/person.jpg" },
|
||
"script": "Your narration here.",
|
||
"voice_id": "<voice-id>"
|
||
}'
|
||
```
|
||
|
||
Common optional fields: `title`, `resolution` (`4k`/`1080p`/`720p`),
|
||
`aspect_ratio`, `remove_background`, `background`, `voice_settings`,
|
||
`motion_prompt` + `expressiveness` (photo-avatar animation), and
|
||
`callback_url`/`callback_id` for webhooks. Don't hardcode these from memory: the
|
||
CLI self-documents the full, current body with
|
||
`heygen video create --request-schema` (a discriminated union keyed on `type`),
|
||
so read the schema rather than trusting a stale field list. For a still you'll
|
||
reuse across many scripts, create a reusable **Photo Avatar** once instead
|
||
(`heygen avatar create`). Ledger the result with
|
||
`resolve --from <downloaded.mp4> --type video`. Docs:
|
||
<https://developers.heygen.com/image-to-video>.
|
||
|
||
## HEVC / H.265 sources
|
||
|
||
HEVC/H.265 sources need no conversion for **render** (FFmpeg pre-decodes all
|
||
input video) or for **preview** (auto-proxy transcodes and caches an H.264
|
||
copy on first use, disable with `--no-proxy` or `media.autoProxy: false` in
|
||
hyperframes.json). A manual H.264 proxy via `ffmpeg -i in.mp4 -c:v libx264
|
||
-crf 18 proxy.mp4`, registered with `resolve --from`, remains available for
|
||
edge cases (e.g. auto-proxy disabled, or ffmpeg unavailable at preview time).
|