mirror of
https://github.com/heygen-com/hyperframes.git
synced 2026-09-01 19:42:03 +00:00
* docs(audio): document the audio effects system The effects feature had no documentation at all — the only mentions anywhere in docs/ were changelog entries. This adds the three pages the color-grading feature already has, one per audience, and closes a gap in the existing audio guide. - prompting/audio-effects — a new Level 5 chapter on asking for a mix in symptoms rather than in filters, with the voiceover carve as the headline, the level-before-depth check, groups, and the three requests that have no honest answer (de-essing, noise removal, tone matching). - studio/audio-effects — the rack, presets by symptom, the carve module, groups, the deliberately asymmetric mute/solo, automation lanes, and troubleshooting. - reference/audio-effects — the contract: all four attributes, every effect and parameter range, why some parameters cannot be automated, the 19 presets, the five jobs and five one-knob profiles, carve semantics, the group model, the render bus, preview/render parity, and the three lint rules. Also points the existing "duck music under important speech" advice in guides/voice-and-audio at the carve, which does that properly and was never named there. Facts verified against source rather than the shipped skill prose: the panel section is "Audio FX" (the neighbouring "Effects" section is visual effects, and an earlier draft of this page named the wrong one), MAX_AUTOMATION_POINTS is 512, fromPreset carries a preset id rather than a boolean, the leveller targets the track's own 80th percentile, group mute reaches the render while solo never leaves Studio, and the CLI carve currently writes clip ids. No screenshots in this pass, and no placeholders for them either — the pages carry concrete markup and tables instead. Screenshots of the rack, the carve module, and a group row would each earn their place later. mint validate and mint broken-links both pass. * docs(audio): give audio its own Studio group, split by task Review feedback: the audio page did not belong in Studio / Edit. That group is already one task per page — canvas, timeline, animation, captions — and a single page covering the rack, the carve, groups, and automation lanes was four tasks bundled together and dropped in beside them. Studio now has an Audio group holding four task pages: - Effects and presets — the rack, presets by symptom, adding single effects in a working order, the one-knob controls, Even Out Levels - Voiceover carve — its own page, because it is the feature people come for - Groups, mute, and solo — including why mute reaches the export and solo cannot - Automation lanes — drawing envelopes, the shape menu, and which parameters cannot move at all Not a new top-level tab: the tabs here are audience-scoped (Guides, Studio, Catalog, Developers), so a feature tab would be the only one of its kind and would strand the prompting chapter out of its Level 5 sequence and the reference page out of Developers. Repointed the deep link in guides/voice-and-audio at the carve's own page. Lane interactions verified in source before documenting: the right-click menu offers Ramp up, Ramp down, Swell, Dip, and Simplify (which needs three points). Deliberately not documented: "clicking a lane label reveals it in the rack", which is not on main. mint validate and mint broken-links both pass. * docs(audio): fix the four review blockers All four verified in source before fixing; the review was right on every count. **The registry was incomplete.** `pitchshift` ships at `audioFx.ts:509-535` (`semitones` −12–12, `mix` 0–1, worklet-backed so neither automatable) and was missing entirely — sixteen effects, not fifteen. It also joins the worklet list, so five effects expose no automatable parameters rather than four. **Three presets were missing.** `chipmunk`, `giant`, and `monster` ship at `audioFxPresets.ts:336-357`, all built on `pitchshift`. Twenty-two presets, and Character holds ten. Fixed in both the reference table and the Studio list. **The copyable markup contradicted the warning above it.** The page said `carve.mjs` only finds double-quoted attributes and then gave three single-quoted examples — copying the chain example would make a later carve miss the existing chain and overwrite it. All three are now double-quoted with `"`, each followed by its unescaped reading so it stays legible. **The attribute table over-claimed.** `data-audio-group` is a plain id, not JSON, and is ignored on `<video>`; the other three also live on `<hf-audio-group>` for a group. The table now carries shape and valid host per attribute. **The automation contract was wrong on two axes.** A clip lane's `t` is clip-relative but a GROUP lane's is composition time, because a group has no `data-start` (`webAudioTransport.ts:337-342`, `audioMixer.ts:1311-1344`) — both pages now split the two clocks. And `volume` is not 0–1: the ceiling is `MAX_AUDIO_GAIN`, +12 dB or about 3.981 (`audioGain.ts:8-9`), so a boosting lane is valid and documented. **Current-main drift.** #3416 is merged, so the CLI now records the voices' shared group when it is safe and falls back to clip ids when that group contains the bed or a music/SFX member. Documented, including why neither refusal shows up on the run that writes it, and rebased onto main. mint validate and mint broken-links both pass. * docs(audio): name the real add-menu family, and finish propagating pitchshift Second review round. All three findings were my own incomplete propagation — I corrected the reference for `pitchshift` last round and left the reader-facing pages behind it. **The add-menu family is `Time`, not `Space`.** `propertyPanelFxAddMenu.tsx:22-28` labels the four groups Filters / Dynamics / Non-linear / Time, and the time group holds pitchshift, delay, chorus, phaser, and reverb. The Studio page sent readers looking for a group that does not exist. It is now a table naming the family and its contents, and the reference's "Time — space and width" heading is retitled, since that description stopped covering the family the moment pitch shift joined it. Also from the same file: the menu offers the named jobs in place of a bare `peaking`, because picking `peaking` is picking a machine and leaving the real decision — which range — for afterwards. Worth saying on the task page. **Pitch shift was missing from both no-automation lists** that a reader actually follows — `studio/audio-automation` and, unflagged but the same defect, the prompting chapter. Five worklet effects in all four places now. Called out explicitly on the Studio page, because a rising pitch is exactly the thing someone reaches for a lane to do, and the lane will not report that it cannot. **Narrowed the group-metadata sentence.** "The other three are JSON, and on a group they live on `<hf-audio-group>`" swept in `data-fx-carve`, contradicting the table directly above it. Only `data-fx-chain` and `data-automation` are group metadata. mint validate and mint broken-links both pass.
125 lines
5.5 KiB
Plaintext
125 lines
5.5 KiB
Plaintext
---
|
|
title: "Make music sit under narration"
|
|
sidebarTitle: "Voiceover carve"
|
|
description: "Take only the frequencies a voice occupies out of a music bed, so the music keeps its character while speech stays intelligible."
|
|
---
|
|
|
|
Music under narration is the most common audio problem in a video, and the
|
|
obvious fix is the wrong one.
|
|
|
|
Turning the whole music track down works, and it costs the music all of its
|
|
presence for as long as anyone is talking — the bed goes limp through the entire
|
|
voiceover. But a voice does not need the whole spectrum. It needs the few bands it
|
|
actually occupies. A **carve** takes only those out of the music, so the bed keeps
|
|
its low end and its top: still music, while the voice stays intelligible.
|
|
|
|
## When to use it
|
|
|
|
Whenever a music bed plays under speech. It is not a polish step to reach if
|
|
there is time — place both tracks, carve, and listen.
|
|
|
|
Skip it only when there is no narration for the music to sit under: a music video,
|
|
a title card, a montage cut to the track.
|
|
|
|
## Carve a bed
|
|
|
|
The carve is a module at the top of the **music** track's rack, not a separate
|
|
tool. Select the music clip, open its **Audio FX** section, and the module asks
|
|
you to pick the voices this bed should make room for.
|
|
|
|
A bed with exactly one candidate voice above it is **already carved**, at the
|
|
default strength, because that is what a bed under narration wants. Several
|
|
candidates leaves the picker waiting rather than guessing which one is the voice.
|
|
|
|
<Warning>
|
|
The carve belongs on the **music**, never on the voice. It names the tracks it
|
|
makes room for, the same way a sidechain compressor does — you select the track
|
|
that gets quieter and pick what makes it quieter. A voice carved against itself
|
|
is a bug, not a subtle mix choice.
|
|
</Warning>
|
|
|
|
## Set the strength
|
|
|
|
One control does the whole job. It derives how deep to cut, how many bands, how
|
|
wide they are, how far the level may drop, and how far under the voice to aim —
|
|
because those move together in any real mix.
|
|
|
|
| Strength | What you get |
|
|
| --- | --- |
|
|
| Default (0.25) | A 6 dB dip in three bands — audible without sounding like a hole |
|
|
| Around 0.5 | The dip reaches 10 dB, where a carve starts being heard as an effect |
|
|
| Higher | Deliberate territory for a loud bed under a quiet voice |
|
|
| 0 | Spectral only — one band, no level matching at all |
|
|
|
|
**Switch the carve off** keeps the settings and stops the carve, which is
|
|
different from clearing it: a bed with one candidate voice would otherwise be
|
|
carved again by default the moment it was re-evaluated.
|
|
|
|
## What it writes
|
|
|
|
Ordinary effects. Open the rack after carving and you will find a few `peaking`
|
|
filters and a `gain` stage — nothing hidden, nothing proprietary.
|
|
|
|
Two things follow from that:
|
|
|
|
- **The envelopes are editable.** The carve always follows the speech, so every
|
|
value becomes an automation lane you can adjust by hand afterwards. Silence
|
|
leaves the music alone; a loud passage pushes the carve to full depth. There is
|
|
no fixed-depth mode, because a fixed dip thins the bed through every pause.
|
|
- **Re-carving is safe.** The nodes a carve writes are tagged as its own, so
|
|
running it again at a new strength replaces exactly those and leaves every
|
|
effect and lane you built by hand where it was.
|
|
|
|
Inside a carved bed the signal runs through the dips first, then the level match,
|
|
then anything you added yourself — which is why a limiter you place stays the last
|
|
ceiling.
|
|
|
|
## Point it at a group, not at clips
|
|
|
|
If narration arrives as several clips, [group them](/studio/audio-groups) and
|
|
carve against the group.
|
|
|
|
A carve pointed at individual clips has to list every one of them, and it stays
|
|
right only until the next edit: add a fourth narration clip and it plays outside
|
|
the carve's awareness, so the music silently fails to duck under it. A carve
|
|
pointed at a group resolves membership every time the analysis runs, so a clip
|
|
added later is covered without touching the carve.
|
|
|
|
Studio does this for you when you pick a second ungrouped voice — it creates a
|
|
group behind them and points the carve at that instead.
|
|
|
|
## What should happen
|
|
|
|
- The voice is legible without the bed sounding hollowed out.
|
|
- The music comes **back up between phrases** rather than staying flat.
|
|
- The bed keeps its bass and its top end. A carve that removed those would just
|
|
be a filter.
|
|
|
|
## Common problems
|
|
|
|
**The music sounds notched or hollow.** Strength is too high. Come back toward
|
|
the default. This is the one failure mode with an obvious sound.
|
|
|
|
**The voice is still buried.** A carve cannot fix a bed that is simply louder than
|
|
the voice — spectral room does not solve a level problem. Compare the two tracks'
|
|
actual loudness first. Machine-generated speech commonly arrives far below a
|
|
mastered music track, and that gap is arithmetic rather than taste.
|
|
|
|
**The bed starts out already ducked.** A lane holds its first value backwards to
|
|
the start of its clip. A bed that begins before the voice needs an explicit "no
|
|
cut" point at its start.
|
|
|
|
**The music pumps.** The level envelope releases slowly on purpose, because music
|
|
that snaps back the instant a word ends sounds like a machine doing it. If it
|
|
still pumps, the gaps between phrases are shorter than the release — lower the
|
|
strength rather than fighting the envelope.
|
|
|
|
**Nothing was carved.** The picker leaves itself alone when several tracks could
|
|
be the voice. Name the voice explicitly.
|
|
|
|
## Related topics
|
|
|
|
- [Mix audio and apply effects](/studio/audio-effects)
|
|
- [Group tracks, mute, and solo](/studio/audio-groups)
|
|
- [Audio effects implementation](/reference/audio-effects#voiceover-carve)
|