Files
hyperframes/docs/studio/voiceover-carve.mdx
Vance Ingalls 0eca7b1e0a docs(audio): document the audio effects system (#3420)
* docs(audio): document the audio effects system

The effects feature had no documentation at all — the only mentions anywhere in
docs/ were changelog entries. This adds the three pages the color-grading feature
already has, one per audience, and closes a gap in the existing audio guide.

- prompting/audio-effects — a new Level 5 chapter on asking for a mix in
  symptoms rather than in filters, with the voiceover carve as the headline, the
  level-before-depth check, groups, and the three requests that have no honest
  answer (de-essing, noise removal, tone matching).
- studio/audio-effects — the rack, presets by symptom, the carve module, groups,
  the deliberately asymmetric mute/solo, automation lanes, and troubleshooting.
- reference/audio-effects — the contract: all four attributes, every effect and
  parameter range, why some parameters cannot be automated, the 19 presets, the
  five jobs and five one-knob profiles, carve semantics, the group model, the
  render bus, preview/render parity, and the three lint rules.

Also points the existing "duck music under important speech" advice in
guides/voice-and-audio at the carve, which does that properly and was never
named there.

Facts verified against source rather than the shipped skill prose: the panel
section is "Audio FX" (the neighbouring "Effects" section is visual effects, and
an earlier draft of this page named the wrong one), MAX_AUTOMATION_POINTS is 512,
fromPreset carries a preset id rather than a boolean, the leveller targets the
track's own 80th percentile, group mute reaches the render while solo never
leaves Studio, and the CLI carve currently writes clip ids.

No screenshots in this pass, and no placeholders for them either — the pages
carry concrete markup and tables instead. Screenshots of the rack, the carve
module, and a group row would each earn their place later.

mint validate and mint broken-links both pass.

* docs(audio): give audio its own Studio group, split by task

Review feedback: the audio page did not belong in Studio / Edit. That group is
already one task per page — canvas, timeline, animation, captions — and a single
page covering the rack, the carve, groups, and automation lanes was four tasks
bundled together and dropped in beside them.

Studio now has an Audio group holding four task pages:

- Effects and presets — the rack, presets by symptom, adding single effects in a
  working order, the one-knob controls, Even Out Levels
- Voiceover carve — its own page, because it is the feature people come for
- Groups, mute, and solo — including why mute reaches the export and solo cannot
- Automation lanes — drawing envelopes, the shape menu, and which parameters
  cannot move at all

Not a new top-level tab: the tabs here are audience-scoped (Guides, Studio,
Catalog, Developers), so a feature tab would be the only one of its kind and
would strand the prompting chapter out of its Level 5 sequence and the reference
page out of Developers.

Repointed the deep link in guides/voice-and-audio at the carve's own page.

Lane interactions verified in source before documenting: the right-click menu
offers Ramp up, Ramp down, Swell, Dip, and Simplify (which needs three points).
Deliberately not documented: "clicking a lane label reveals it in the rack",
which is not on main.

mint validate and mint broken-links both pass.

* docs(audio): fix the four review blockers

All four verified in source before fixing; the review was right on every count.

**The registry was incomplete.** `pitchshift` ships at `audioFx.ts:509-535`
(`semitones` −12–12, `mix` 0–1, worklet-backed so neither automatable) and was
missing entirely — sixteen effects, not fifteen. It also joins the worklet list,
so five effects expose no automatable parameters rather than four.

**Three presets were missing.** `chipmunk`, `giant`, and `monster` ship at
`audioFxPresets.ts:336-357`, all built on `pitchshift`. Twenty-two presets, and
Character holds ten. Fixed in both the reference table and the Studio list.

**The copyable markup contradicted the warning above it.** The page said
`carve.mjs` only finds double-quoted attributes and then gave three
single-quoted examples — copying the chain example would make a later carve miss
the existing chain and overwrite it. All three are now double-quoted with
`"`, each followed by its unescaped reading so it stays legible.

**The attribute table over-claimed.** `data-audio-group` is a plain id, not JSON,
and is ignored on `<video>`; the other three also live on `<hf-audio-group>` for
a group. The table now carries shape and valid host per attribute.

**The automation contract was wrong on two axes.** A clip lane's `t` is
clip-relative but a GROUP lane's is composition time, because a group has no
`data-start` (`webAudioTransport.ts:337-342`, `audioMixer.ts:1311-1344`) — both
pages now split the two clocks. And `volume` is not 0–1: the ceiling is
`MAX_AUDIO_GAIN`, +12 dB or about 3.981 (`audioGain.ts:8-9`), so a boosting lane
is valid and documented.

**Current-main drift.** #3416 is merged, so the CLI now records the voices'
shared group when it is safe and falls back to clip ids when that group contains
the bed or a music/SFX member. Documented, including why neither refusal shows
up on the run that writes it, and rebased onto main.

mint validate and mint broken-links both pass.

* docs(audio): name the real add-menu family, and finish propagating pitchshift

Second review round. All three findings were my own incomplete propagation — I
corrected the reference for `pitchshift` last round and left the reader-facing
pages behind it.

**The add-menu family is `Time`, not `Space`.** `propertyPanelFxAddMenu.tsx:22-28`
labels the four groups Filters / Dynamics / Non-linear / Time, and the time group
holds pitchshift, delay, chorus, phaser, and reverb. The Studio page sent readers
looking for a group that does not exist. It is now a table naming the family and
its contents, and the reference's "Time — space and width" heading is retitled,
since that description stopped covering the family the moment pitch shift joined
it.

Also from the same file: the menu offers the named jobs in place of a bare
`peaking`, because picking `peaking` is picking a machine and leaving the real
decision — which range — for afterwards. Worth saying on the task page.

**Pitch shift was missing from both no-automation lists** that a reader actually
follows — `studio/audio-automation` and, unflagged but the same defect, the
prompting chapter. Five worklet effects in all four places now. Called out
explicitly on the Studio page, because a rising pitch is exactly the thing
someone reaches for a lane to do, and the lane will not report that it cannot.

**Narrowed the group-metadata sentence.** "The other three are JSON, and on a
group they live on `<hf-audio-group>`" swept in `data-fx-carve`, contradicting the
table directly above it. Only `data-fx-chain` and `data-automation` are group
metadata.

mint validate and mint broken-links both pass.
2026-08-22 05:49:35 -07:00

125 lines
5.5 KiB
Plaintext

---
title: "Make music sit under narration"
sidebarTitle: "Voiceover carve"
description: "Take only the frequencies a voice occupies out of a music bed, so the music keeps its character while speech stays intelligible."
---
Music under narration is the most common audio problem in a video, and the
obvious fix is the wrong one.
Turning the whole music track down works, and it costs the music all of its
presence for as long as anyone is talking — the bed goes limp through the entire
voiceover. But a voice does not need the whole spectrum. It needs the few bands it
actually occupies. A **carve** takes only those out of the music, so the bed keeps
its low end and its top: still music, while the voice stays intelligible.
## When to use it
Whenever a music bed plays under speech. It is not a polish step to reach if
there is time — place both tracks, carve, and listen.
Skip it only when there is no narration for the music to sit under: a music video,
a title card, a montage cut to the track.
## Carve a bed
The carve is a module at the top of the **music** track's rack, not a separate
tool. Select the music clip, open its **Audio FX** section, and the module asks
you to pick the voices this bed should make room for.
A bed with exactly one candidate voice above it is **already carved**, at the
default strength, because that is what a bed under narration wants. Several
candidates leaves the picker waiting rather than guessing which one is the voice.
<Warning>
The carve belongs on the **music**, never on the voice. It names the tracks it
makes room for, the same way a sidechain compressor does — you select the track
that gets quieter and pick what makes it quieter. A voice carved against itself
is a bug, not a subtle mix choice.
</Warning>
## Set the strength
One control does the whole job. It derives how deep to cut, how many bands, how
wide they are, how far the level may drop, and how far under the voice to aim —
because those move together in any real mix.
| Strength | What you get |
| --- | --- |
| Default (0.25) | A 6 dB dip in three bands — audible without sounding like a hole |
| Around 0.5 | The dip reaches 10 dB, where a carve starts being heard as an effect |
| Higher | Deliberate territory for a loud bed under a quiet voice |
| 0 | Spectral only — one band, no level matching at all |
**Switch the carve off** keeps the settings and stops the carve, which is
different from clearing it: a bed with one candidate voice would otherwise be
carved again by default the moment it was re-evaluated.
## What it writes
Ordinary effects. Open the rack after carving and you will find a few `peaking`
filters and a `gain` stage — nothing hidden, nothing proprietary.
Two things follow from that:
- **The envelopes are editable.** The carve always follows the speech, so every
value becomes an automation lane you can adjust by hand afterwards. Silence
leaves the music alone; a loud passage pushes the carve to full depth. There is
no fixed-depth mode, because a fixed dip thins the bed through every pause.
- **Re-carving is safe.** The nodes a carve writes are tagged as its own, so
running it again at a new strength replaces exactly those and leaves every
effect and lane you built by hand where it was.
Inside a carved bed the signal runs through the dips first, then the level match,
then anything you added yourself — which is why a limiter you place stays the last
ceiling.
## Point it at a group, not at clips
If narration arrives as several clips, [group them](/studio/audio-groups) and
carve against the group.
A carve pointed at individual clips has to list every one of them, and it stays
right only until the next edit: add a fourth narration clip and it plays outside
the carve's awareness, so the music silently fails to duck under it. A carve
pointed at a group resolves membership every time the analysis runs, so a clip
added later is covered without touching the carve.
Studio does this for you when you pick a second ungrouped voice — it creates a
group behind them and points the carve at that instead.
## What should happen
- The voice is legible without the bed sounding hollowed out.
- The music comes **back up between phrases** rather than staying flat.
- The bed keeps its bass and its top end. A carve that removed those would just
be a filter.
## Common problems
**The music sounds notched or hollow.** Strength is too high. Come back toward
the default. This is the one failure mode with an obvious sound.
**The voice is still buried.** A carve cannot fix a bed that is simply louder than
the voice — spectral room does not solve a level problem. Compare the two tracks'
actual loudness first. Machine-generated speech commonly arrives far below a
mastered music track, and that gap is arithmetic rather than taste.
**The bed starts out already ducked.** A lane holds its first value backwards to
the start of its clip. A bed that begins before the voice needs an explicit "no
cut" point at its start.
**The music pumps.** The level envelope releases slowly on purpose, because music
that snaps back the instant a word ends sounds like a machine doing it. If it
still pumps, the gaps between phrases are shorter than the release — lower the
strength rather than fighting the envelope.
**Nothing was carved.** The picker leaves itself alone when several tracks could
be the voice. Name the voice explicitly.
## Related topics
- [Mix audio and apply effects](/studio/audio-effects)
- [Group tracks, mute, and solo](/studio/audio-groups)
- [Audio effects implementation](/reference/audio-effects#voiceover-carve)