Files
hyperframes/skills/media-use/SKILL.md
T
Miguel ÁngelandClaude Opus 4.8 57b3c78987 feat(media-use): color grading — grade/lut resolve, smart-grade, grade-compare + compare (#2041)
* feat(media-use): color grading — grade/lut resolve, smart-grade, grade-compare CLI

Add color grading to media-use as first-class resolve types plus a faithful
comparison command. All local, offline, deterministic — no model, no GPU.

- resolve -t grade / -t lut: produce a data-color-grading block (or a frozen
  .cube). Look cascade: core preset (no file) -> bundled .cube library ->
  parametric buildCube. Emitted .cube is Rec.709 and validated against core's
  colorLuts constraints (LUT_3D_SIZE <= 64) before it is frozen.
- smart grade (grade --for <media>): ffmpeg signalstats -> adjust suggestion
  (exposure / contrast / white balance), surfaced with the measured evidence on
  stderr as a starting point; never auto-applied.
- hyperframes grade-compare: renders N candidate grades onto a reference frame
  through the real runtime shader into one labeled comparison PNG, so an agent
  picks a look without opening Studio. Prepends an "original" baseline cell by
  default (--no-baseline to omit). Shares the headless-capture pipeline with
  snapshot via capture/captureCompositionFrame.
- media-use SKILL: proactive "media opportunity pass" guidance (grounded
  signal -> offer, ask once, surface don't mutate).

Verified: media-use 116/116, grade-compare 7/7, snapshot 9/9, lint + format
clean, full build green, comparison renders end to end.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

* test(cli): narrow grade-compare baseline assertion off unknown-typed grading

Assert the whole cell via toEqual instead of reaching into .grading.preset /
.grading.lut on the unknown-typed field, keeping the test typecheck-clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

* feat(media-use): agent-authored LUTs via --params + validate --from cube; never-read-.cube guardrail

- resolve -t lut / -t grade --params '<json>': build a parametric .cube from
  explicit params (bypassing the intent cascade), validate, and freeze in one
  step. --intent becomes the optional description. Lets an agent commit a look
  it computed itself.
- --from <file.cube> now validates the ingested LUT for lut/grade types and
  rejects an invalid/oversized cube (no partial write) — the escape hatch for a
  LUT the agent generated with its own code.
- SKILL.md: hard rule to never read a .cube body into context (~size^3 lines,
  zero legible signal) — inspect via grade-compare (see it) or cube-validate
  (ok/size), read the manifest description for meaning; plus both authoring
  paths and the parametric-vs-film-stock ceiling note.

Verified: media-use 116/116, lint + format clean; smokes — --params builds a
valid frozen cube, grade --params returns a lut block, bad JSON and an oversized
--from cube are both rejected with no stray file.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

* fix(cli): grade-compare validates referenced LUTs, warns on no-op cells, caps candidates

Bug-bash follow-ups — grade-compare silently accepted bad input:

- Validate LUT *content*, not just existence: each referenced .cube is parsed
  with core's parseCubeLut (now exported from @hyperframes/core) and rejected
  with a per-cell error ("LUT for \"<label>\" is not a valid .cube: ..."). A
  file that exists but isn't a valid cube no longer renders a silent no-op cell.
- Warn on inactive cells: a grading that normalizes to inactive (e.g. a
  malformed {lut:12345}) emits a stderr warning naming the cell; the
  auto-prepended "original" baseline is intentionally inactive and stays silent.
  stdout remains valid JSON.
- Cap candidates at 16 (excluding baseline): over-cap input renders the first N
  and reports {truncated:true, total:M} on stdout + a stderr note — no silent
  drop, no unbounded giant sheet.

Verified: grade-compare 10/10; non-cube LUT → clear error; {lut:12345} → warning
+ ok; 20 cells → cells=17 truncated total=20; valid runs unchanged. Lint/format
clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

* feat(cli): general `hyperframes compare` visual-variant primitive

Generalize grade-compare's "render N variants → one labeled sheet → the agent
looks and picks" loop into a standalone command that works on ANY variation
(font, layout, motion, grade, whole compositions) — the tool never needs to
know what differs.

- `hyperframes compare <path...> [--at <sec>] [--labels a,b,c] [--out] [--cols]
  [--json]`: renders each agent-authored composition variant through the real
  runtime (captureCompositionFrame) and stitches one labeled comparison sheet +
  JSON ({ok, sheet, rendered, variants, truncated?/total?}). 2+ paths required;
  caps at 16 with loud truncation. It presents, it does not judge — choosing is
  the caller's job.
- Factored the shared "render a labeled set → contact sheet" path so compare,
  grade-compare, and snapshot all sit on it (no duplication). grade-compare is
  now the first color-specific specialization of this primitive.
- New pathArgs util + contactSheet test; hyperframes-cli SKILL documents compare
  as the agent's "see your own renders and choose" primitive.

Verified: 26/26 across compare + grade-compare + snapshot + contactSheet (no
regressions); compare renders 3 variants into one visibly-distinct labeled
sheet; 2+-path error path clean; lint/format clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

* fix(ci): green the skills CI — skip ffmpeg tests when absent, oxfmt markdown

The "Test: skills" CI job runs bare `node --test` with no ffmpeg on PATH (by
design — skills tests are meant to be node-builtin-only). The grade-analyzer +
smart-grade tests shell to ffmpeg and were failing there with ENOENT. Guard
them to skip when ffmpeg isn't on PATH; they still run locally / where it is.

Also oxfmt README.md + hyperframes/media-use SKILL.md (the whole-repo
`oxfmt --check .` Format job caught markdown left unformatted by the rebase
conflict resolution).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

* fix(ci): skip core-conformance test when tsx is unavailable

The "Test: skills" CI job installs no deps, so the normalizeHfColorGrading
conformance test (which imports core's TS via `node --import tsx`) failed there.
Guard it to skip when tsx can't resolve; runs locally / in the deps-installed
Test job. Completes the skills-CI greening (the ffmpeg guards handled the rest).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

* fix(cli): escape grade-compare src double-quotes (CodeQL XSS) + Windows-safe compare test

- grade-compare built `<img src="...">` (double-quoted) with the single-quote
  escaper, leaving `"` unescaped — a `"` in the frame path could break out
  (CodeQL: incomplete HTML attribute sanitization). Use escapeXml for src.
- compare label test hard-coded POSIX paths that can't match on Windows; assert
  the derived labels (the subject); path resolution is covered elsewhere.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

* refactor(media-use): generate LUT library from params (drop committed .cube files)

The 3 bundled .cube files were 733 lines each (2,199 total) and were themselves
buildCube output — pure repo bloat. Replace with compact per-look params in
luts/index.json, generated on resolve; add an optional `url` for future scanned
LUTs to be CDN-hosted + downloaded on demand (freezeUrl) instead of committed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

* feat(media-use): serve library LUTs from CDN on-demand (static.heygen.ai/luts), params fallback

Looks now carry a CDN `url` (hosted at s3://heygen-public/luts → static.heygen.ai/luts/<id>.cube);
resolve downloads + validates + freezes on demand, like bgm/image. `params` stays
as the deterministic offline fallback (--local-only, or if the download fails), so
resolution is never blocked on the network. Provider prefers url, falls back to params.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

* fix(media-use): address #2041 review — atomic LUT writes, compare telemetry, follow-ups

- Atomic .cube writes: library provider (url + params) and the parametric
  generator now write to a .tmp path, validate, then rename, so a crash can
  never orphan an invalid .cube at the final path (was validate-after-write).
- track("media_use_resolve") now emits provenance.via (url/params-fallback/params).
- grade-compare + compare: --timeout flag (was hardcoded 5000) and a
  media_use_compare event (cells, truncated, total, render_ready_timed_out);
  openSettledCompositionPage now surfaces the render-ready timeout.
- compare staging skips node_modules/.git; --for gets an upfront existence check.
- Rec.709 luma comment; HYPERFRAMES_ANALYZE_TIMEOUT_MS override; measured note
  uses basename; LUT s3 hosting moved from index.json into luts/README.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-08 22:20:16 -04:00

26 KiB
Raw Blame History

name, description
name description
media-use Agent Media OS, the single skill for every media need in a HyperFrames project. Resolve BGM, SFX, image, icon, brand logo, voice, color grade, or LUT into a frozen local file or paste-ready block + ledger record (one verb, `resolve`); generate via TTS / music / image models when the catalog misses; produce voiceover, transcription, captions, and background removal through one shared audio engine; operate on media (cut / reframe / transform); and reuse assets across projects. Keeps search noise on disk, hands the agent one path or block. Use for any audio, image, icon, logo, voiceover, caption, color-grading, or media-asset need.

media-use

The media OS for HyperFrames: resolve · generate · operate · remember, every media type, one skill, zero context noise.

What it owns (the gaps HyperFrames leaves)

HyperFrames owns media playback; media-use owns everything else. Each row is enforced by scripts/lib/coverage.test.mjs so the claim can't rot.

HyperFrames gap media-use owns it via
Audio-only, no image/icon resolve --type image|icon (heygen asset search)
No third-party brand logos resolve --type logo (svgl → simple-icons → GitHub org avatar → domain favicon)
No voice / audio generation resolve --type voice + the audio engine (audio/scripts/audio.mjs)
Scattered/duplicated audio engine one consolidated engine under audio/ (hyperframes-media retired)
No agent media-ops (cut/reframe/transform) references/operations.md + resolve --from to register outputs
No transcript-driven cutting scripts/transcript-cut.mjs compiles word-timestamp edits into cut lists
No auto-duck / publish loudness scripts/audio-duck.mjs + references/operations.md loudnorm/sidechain recipes
No cross-project memory global content-addressed cache + auto-promote (~/.media)
No color-grade authoring resolve --type grade emits a paste-ready data-color-grading block; resolve --type lut freezes validated .cube files
No image generation RAM-graded local mflux (FLUX) via scripts/lib/mflux-provider.mjs, codex image_gen upsell (scripts/lib/codex-provider.mjs)
No video generation spec-gated local LTX (videogen in scripts/lib/local-models.mjs); heygen video create avatar upsell
Weak local-model defaults free-usage HeyGen first (TTS, bg-removal) via the heygen CLI; local open-source only as an opt-out fallback (scripts/lib/local-run.mjs)

When to use

Call resolve whenever a composition needs media: background music, sound effects, images, icons, brand logos, voice, a color grade, or a LUT. For voiceover / TTS, music, SFX, and caption timing, use the audio engine (below); background removal is delegated to the hyperframes CLI; transcription defaults to Parakeet (better than whisper.cpp: 6.05% vs 7.44% WER, 5-10x faster) via scripts/transcribe.mjs, with whisper.cpp auto-fallback (see references/operations.md). For cutting / reframing / transforming existing media, see references/operations.md. media-use searches the HeyGen catalog first for media files, resolves official logos through the logo cascade, uses local deterministic color grading for grade/lut, freezes the best match locally when a file is needed, registers it in a manifest, and hands the agent one line; all search noise stays on disk.

Be proactive — run a media opportunity pass

The human usually can't tell which media would lift the piece. You can. When you build or review a composition, do one grounded scan and then ask once — don't silently add, and don't nag per asset.

Surface an opportunity only when a concrete signal is present:

Signal detected Offer
On-screen text / a script with no voiceover TTS voiceover (audio engine)
Emoji or a <div> styled as an icon resolve real icons
Image that is a placeholder, tiny, or upscaled-looking a better image (and/or upscale — see references/operations.md)
Hard scene cuts / transitions with no sound transition sfx
A piece over ~10s with no music bed bgm
Footage that reads under/over-exposed or color-cast a corrective grade (analyze with grade --for, preview with hyperframes grade-compare)

Rules that keep this a help, not nagware:

  • Grounded, not generic. No signal → no suggestion. Never open with "want better images?".
  • Opinionated + concrete. Propose the specific fix ("add a VO from your script, swap 3 emoji for real icons, replace the 400×400 hero, whooshes on the 4 cuts"), with defaults chosen — the human just approves all / some / none.
  • Once per project. One consolidated ask, top few highest-value items. Respect "leave it" and don't re-raise.
  • Surface, never silently mutate. Color grades especially: propose and preview, never auto-apply — a gray-world "correction" ruins an intentional sunset or neon look.

Resolve

node <SKILL_DIR>/scripts/resolve.mjs --type <type> --intent "<description>" --project <dir>

Returns one line: resolved <id> → <path> (<type>, <metadata>)

Types

Type What it finds Provider / cascade
bgm Background music HeyGen audio catalog (10k+ tracks)
sfx Sound effects Bundled 19-file library + HeyGen catalog
image Photos, backgrounds HeyGen asset search (75k+ vectors)
icon Icons, symbols HeyGen asset search (type=icon)
logo Official brand marks svgl → simple-icons → GitHub org avatar → domain favicon
voice TTS voiceover Local Kokoro (free); HeyGen TTS upsell
grade HyperFrames color-grading blocks Core preset → look index params/CDN LUT → deterministic cube
lut Reusable .cube LUT files Look index params/CDN LUT → deterministic cube

Examples

# Background music
node <SKILL_DIR>/scripts/resolve.mjs --type bgm --intent "upbeat tech launch" --project .
# → resolved bgm_001 → .media/audio/bgm/bgm_001.mp3 (bgm, 25s)

# Sound effect
node <SKILL_DIR>/scripts/resolve.mjs --type sfx --intent "whoosh" --project .
# → resolved sfx_001 → .media/audio/sfx/sfx_001.mp3 (sfx, 0.57s)

# Image
node <SKILL_DIR>/scripts/resolve.mjs --type image --intent "gradient tech background" --project .
# → resolved image_001 → .media/images/image_001.jpg (image)

# Icon
node <SKILL_DIR>/scripts/resolve.mjs --type icon --intent "rocket" --project .
# → resolved icon_001 → .media/images/icon_001.png (icon, transparent)

# Brand logo (official mark — never redrawn by hand)
node <SKILL_DIR>/scripts/resolve.mjs --type logo --entity linkedin --intent "LinkedIn logo" --project .
# → resolved logo_001 → .media/images/logo_001.svg (logo, official mark)

# Color grade block
node <SKILL_DIR>/scripts/resolve.mjs --type grade --intent "warm daylight" --project . --json
# → {"ok":true,"preset":"warm-daylight","grading":{"preset":"warm-daylight","intensity":1},...}

# LUT file
node <SKILL_DIR>/scripts/resolve.mjs --type lut --intent "teal orange blockbuster" --project .
# → resolved lut_001 → .media/luts/lut_001.cube (lut)

Flags

Flag Description
--type, -t Media type: bgm, sfx, image, icon, logo, voice, grade, lut
--intent, -i What you need (natural language)
--entity, -e Entity name for cache matching (optional)
--project, -p Project directory (default: .)
--candidates List reusable assets (project + global cache) for --type; no download, no mutation
--reuse <sha> Import a specific global-cache asset (by content sha/prefix, from --candidates)
--from Freeze a local file or direct public URL (ingest)
--for Analyze a local image/video and add measured adjust suggestions (grade only)
--local-only Offline: skip every network provider (cache + local only)
--provider Force one generator (e.g. codex, mflux, kokoro, heygen)
--adopt Bulk-import existing assets/ into manifest
--json Output JSON instead of one-line result

Reuse before you resolve

Before resolving bgm/sfx/image/icon/logo/grade/lut, check what already exists and reuse it when it fits. media-use does not semantically match for you — you are the judge. It surfaces candidates; you decide.

node <SKILL_DIR>/scripts/resolve.mjs --type bgm --intent "upbeat tech launch" --candidates --project .
#   [project] upbeat tech launch (25s, heygen.audio.sounds)
#           .media/audio/bgm/bgm_001.wav
#   [global]  energetic tech intro (22s, heygen.audio.sounds)
#           --reuse 06e052c075fd2b80

Read the list and judge semantic fit yourself — "upbeat tech launch" ≈ "energetic tech intro" is a call only you can make from the descriptions. Then:

  • A project candidate fits → just reference its path in your composition. Nothing else to run.
  • A global candidate fitsresolve --type bgm --reuse <sha> copies it into this project (self-contained render) and records it.
  • Nothing fits → resolve fresh (--type ... --intent ...).

Trust guardrail — when unsure, resolve fresh. A redundant download is cheap; shipping the wrong asset is not. Judge fit from description + prompt + type + duration/dims. For brand/entity assets, reuse a global candidate only when the entity matches exactly — the global cache aggregates every project you have worked on, so a --candidates list can surface another client's brand mark and its prompt text. Never reuse a cross-project brand asset on a loose match.

The deterministic floor still runs automatically: an identical (case/whitespace-insensitive) repeat auto-reuses with no --candidates step. --candidates is only for the semantic layer above that floor — and a fuzzy match is never auto-applied; reuse is always your explicit call. On a resolve that misses the floor and is about to fetch, media-use prints a one-line stderr hint when similar cached assets exist, pointing you back here.

Color grading

Use grade when you need the actual HyperFrames data-color-grading value to paste onto an <img> or <video>. Core presets and params-backed library looks resolve locally; future CDN-backed library looks require network unless already frozen:

Never cat/read a .cube file into context. A 3D LUT is ~size^3 lines of raw numbers (33^3 ≈ 36k lines at the default size). It bloats context and carries zero human/agent-legible signal. To understand or choose a LUT, use hyperframes grade-compare to see it rendered, or cube-validate.mjs for a one-line {ok,size} check. Read .media/index.md or luts/index.json for the description. Never read the LUT body itself.

node <SKILL_DIR>/scripts/resolve.mjs --type grade --intent "warm daylight" --project . --json

Preset-first output uses the core runtime vocabulary and does not freeze a file:

{
  "preset": "warm-daylight",
  "intensity": 1
}

Paste it as an attribute value after JSON string escaping:

<video
  class="clip"
  src="./media/scene.mp4"
  data-color-grading='{"preset":"warm-daylight","intensity":1}'
></video>

Looks beyond the preset vocabulary freeze a validated .cube under .media/luts/ and return a block that references it:

node <SKILL_DIR>/scripts/resolve.mjs --type grade --intent "teal orange blockbuster" --project . --json
{
  "intensity": 1,
  "lut": { "src": ".media/luts/grade_001.cube", "intensity": 0.85 }
}

Use lut when you only need the reusable .cube file:

node <SKILL_DIR>/scripts/resolve.mjs --type lut --intent "teal orange blockbuster" --project .

For a describable technical look, author an explicit parametric LUT with --params:

node <SKILL_DIR>/scripts/resolve.mjs --type lut --params '{"contrast":0.2,"temperature":-0.3}' --project .
node <SKILL_DIR>/scripts/resolve.mjs --type grade --params '{"exposure":0.2}' --project . --json

For a LUT generated by your own script, ingest it with --from; media-use validates it before registration and rejects invalid or oversized cubes:

node <SKILL_DIR>/scripts/resolve.mjs --type lut --from custom.cube --project .

Parametric math (buildCube) cannot reproduce real film stocks or emulsion looks. Use a CDN-backed scanned .cube entry or ingest a real scanned .cube for those.

For visual selection, list reusable looks with resolve --type grade --candidates, write the promising entries to a grades.json, run hyperframes grade-compare --for <frame> --grades grades.json, then commit the winner with resolve -t grade as the final data-color-grading block.

Smart grade is grade --for <media>. It runs local ffmpeg/ffprobe signalstats, merges a bounded adjust suggestion into the returned block, and prints the measured evidence to stderr. Stdout remains valid JSON under --json; the suggestion is a starting point for the agent to tune, not an automatic neutralization of intentional color.

node <SKILL_DIR>/scripts/resolve.mjs --type grade --intent "warm cinematic" --for ./frame.png --project . --json

Library looks live in luts/index.json. Each entry keeps id, description, tags, and intensity, then supplies either compact params for on-demand buildCube(params) generation or a direct CDN url for future scanned .cube files. Do not commit generated .cube bodies; resolve validates generated or downloaded cubes as it freezes them under .media/luts/.

node skills/media-use/scripts/resolve.mjs --type lut --intent "teal orange blockbuster" --project . --json
node skills/media-use/scripts/lib/cube-validate.mjs .media/luts/lut_001.cube

Providers

media-use holds no keys; every external tool owns its auth. Generation is local-first with a cloud upsell where one helps. resolve spec-checks AVAILABLE RAM and auto-picks the best local model that fits (a RAM-graded ladder, describeModelLadder); the agent can see the ladder and override.

Type Provider (in order)
bgm/sfx heygen catalog (free)
image heygen search, then local mflux (best FLUX for your RAM), then codex image_gen upsell
voice local Kokoro (free, on-device), then heygen tts paid upsell
icon heygen asset search
logo svgl, then simple-icons, then GitHub org avatar, then domain favicon (all free)
grade/lut local core-preset map, params/CDN look index, deterministic buildCube fallback
video (local) local LTX (videogen ladder); heygen video create avatar upsell

Local Kokoro (voice), mflux (image), and LTX (video) run on-device (free, private, offline once cached). Paid/cloud upsells sit behind them: HeyGen TTS for voice, the codex CLI (ChatGPT sub) for a better image, the heygen CLI for avatar video. Cost rule (X4): the agent confirms before an agent-initiated paid call; a user-requested one just runs.

To force a specific generator (e.g. a user says "make this image with codex"), pass --provider codex: it pins resolution to that provider and skips the free-first default. See references/operations.md for the RAM ladders and upsell recipes.

--local-only skips every network provider, including the free HeyGen ones, leaving the project + global cache and any local provider.

How it works

resolve runs an automatic floor, then falls through to fetching:

  1. Check project .media/manifest.jsonl for a prompt match (case- and whitespace-insensitive) — auto-reuse
  2. Scan existing assets/ directory for unregistered files that share a word with the need
  3. Check global cache ~/.media/ for a reusable asset matched on the same normalized prompt — auto-reuse
  4. Search via provider (HeyGen audio catalog, HeyGen asset search), or resolve color locally
  5. Freeze file to .media/<type>/, register in manifest, regenerate index.md, auto-promote to ~/.media/

Steps 1 and 3 are the deterministic floor: they only auto-reuse an exact-normalized match, never a fuzzy one. Semantic reuse ("close enough") is the agent's explicit call via Reuse before you resolve — it never happens automatically. The agent gets back one line; candidates, scores, provenance stay on disk.

Adopt existing projects

Most HyperFrames projects already have assets in assets/. media-use adopts them:

node <SKILL_DIR>/scripts/resolve.mjs --adopt --project .
# → adopted 9 assets from assets/
#   bgm_001 → assets/bgm/mango-fizz.mp3 (bgm, 146.6s)
#   image_001 → assets/images/avatar.jpg (image, 400×400)

ffprobe extracts real duration and dimensions. During resolve, unregistered files in assets/ matching the intent are adopted on the fly.

Reading the inventory

After resolve or adopt, read .media/index.md for the full inventory:

# .media · 4 assets

id         type   dur   dims       path                          description
bgm_001    bgm    25s   -          .media/audio/bgm/bgm_001.mp3  upbeat tech launch
sfx_001    sfx    0.6s  -          .media/audio/sfx/sfx_001.mp3  whoosh
image_001  image  -     1920×1080  .media/images/image_001.jpg   gradient tech background
icon_001   icon   -     200×200    .media/images/icon_001.png    rocket

Cross-project reuse

Assets are cached automatically on resolve. Every resolved/ingested asset is auto-promoted to the global cache at ~/.media/, so subsequent resolves for the same (or near-identical) prompt, in any project, hit the cache with no re-download and no provider call.

For a semantically similar (not identical) need in another project, the exact-match floor won't fire — use Reuse before you resolve: --candidates lists the global assets, and --reuse <sha> imports the one you pick. This is how a track resolved in one project gets reused in the next when the wording differs.

Files

  • .media/manifest.jsonl: machine SSOT, one JSON record per line
  • .media/index.md: agent-readable table (id, type, dur, dims, path, description)
  • ~/.media/: global cross-project reuse cache (content-addressed, SHA-256)

Audio engine: voiceover, music, SFX, captions, transcription

For a full audio pass (TTS voiceover + background music + sound effects in one shot), use the shared engine at audio/scripts/audio.mjs. It takes a neutral audio_request.json and writes audio_meta.json plus assets under .media/audio/{voice,bgm,sfx}:

node <SKILL_DIR>/audio/scripts/audio.mjs --request ./audio_request.json --out ./audio_meta.json
  • Request { provider?, lang?, speed?, lines: [{ id, text, sfx?: [names] }], bgm: { mode?, query?, prompt? } }: id joins each line back to your model; bgm.mode = retrieve | generate | none (omit for auto). --only tts,bgm,sfx runs a subset and merges into an existing --out.
  • Output audio_meta.json (id-keyed): voices[].{path,duration_s,words[]} (word timestamps for captions), sfx[], bgm, total_duration_s.
  • Auto-degrades on one switch: HeyGen credential present → HeyGen TTS + music/SFX retrieval; absent → ElevenLabs/Kokoro TTS, Lyria/MusicGen BGM generation, and the bundled SFX library (no credential needed).
  • If BGM took the generate path (bgm_pending: true), run audio/scripts/wait-bgm.mjs before final render.

Single-shot helpers: audio/scripts/heygen-tts.mjs (one voice file). Transcription / background removal / captions use the hyperframes CLI (transcribe, remove-background), see the per-topic guides in audio/references/ (tts.md, bgm.md, sfx.md, transcribe.md, remove-background.md, captions/).

Operating on media (cut, reframe, transform)

media-use resolves + remembers; for operating on assets see references/operations.md: local-tool recipes (ffmpeg trim/reframe/montage, auto-editor, scenedetect) and the local-vs-HeyGen transform table (background removal, upscale, lipsync, translate). Run the tool, then register the output with resolve --from <output> --type <type> so it joins the ledger + global cache.

CLI tools used (what to run, and how to enable each)

resolve auto-cascades; each provider shells one CLI. Local tools are OPT-IN: if a local tool is absent, resolve degrades gracefully to the free/cloud path, so nothing here is strictly required except ffmpeg/ffprobe. Install a local tool to unlock its free, private, on-device path. media-use holds no keys.

Tool Serves Install
ffmpeg/ffprobe adopt probing, smart-grade signalstats, cut, duck bake, loudnorm system package (brew install ffmpeg)
heygen catalog (bgm/sfx/image/icon), TTS + avatar upsell curl -fsSL https://static.heygen.ai/cli/install.sh | bash then heygen auth login --key <key> (needs >= v0.1.6)
mflux-generate local image gen (FLUX), best-for-RAM uv venv ~/.venvs/mflux && VIRTUAL_ENV=~/.venvs/mflux uv pip install mflux==0.9.6
codex image gen upsell (ChatGPT sub) Codex CLI, logged in via ChatGPT (owns its own auth)
parakeet-mlx local transcription (default ASR, best) uv venv ~/.venvs/parakeet && VIRTUAL_ENV=~/.venvs/parakeet uv pip install parakeet-mlx
ltx-2-mlx local video gen git clone https://github.com/dgrauet/ltx-2-mlx && cd ltx-2-mlx && uv sync --all-extras
npx hyperframes Kokoro TTS (voice), whisper.cpp (transcribe fallback), remove-background bundled with the hyperframes CLI

The RAM-graded local-model shortlist + exact per-tier install/invoke lives in scripts/lib/local-models.mjs (the agent can read describeModelLadder(cap, specs) to see which model fits this machine). Without a tool on PATH, its provider prints a one-line diagnostic to stderr and resolve falls through to the next provider (e.g. no mflux -> codex image upsell; no parakeet-mlx -> whisper.cpp).

heygen asset search is a pre-launch command hidden from heygen --help, but it runs; providers tag requests with the allowlisted X-HeyGen-Client-Source header (v0.1.6+).

Telemetry

resolve and the edit tools (transcribe / transcript-cut / audio-duck) send an anonymous usage event to PostHog (scripts/lib/telemetry.mjs), so we can see which capabilities are actually used. It records only the media TYPE, the resolution SOURCE, and the winning PROVIDER: never the intent text, file names, or paths, and $ip:null so no IP is stored. Best-effort and non-blocking (a resolve never waits on or fails from telemetry).

Opt out with DO_NOT_TRACK=1 or HYPERFRAMES_NO_TELEMETRY=1 (also off in CI and dev). Same public PostHog project key and opt-outs as the hyperframes CLI.