Talk2Camera
Menu
en-GB
video editing

How to duck background music under your voice

A music bed at one constant level fights the voice on top of it. Broadcast mixers ride the fader; ducking does the riding for you.

Music under a talking video fails in a specific, predictable way. You pick a track, set its volume so the intro feels alive, and press export — and the finished video is a fight. Not because the music is bad, but because a constant level is the wrong shape for speech. A voice is not constant: it rises into a point, drops at the end of a sentence, disappears entirely while you breathe. A music bed held at one level is too loud whenever you speak softly and too quiet the moment you stop, so the viewer’s ear is forever renegotiating which sound to follow. Worse, voice and music live in the same place. The instruments that make a track feel warm — guitars, keys, strings, most vocals — occupy the same few octaves as human speech, so the two do not stack, they mask. Turning the whole track down does not solve this; it just moves the fight to a lower volume, and the pauses now feel empty instead of scored. The problem is not the level. It is that the level never changes.

Broadcast solved this decades ago with a hand on a fader. Listen closely to any well-made documentary or radio piece: the music is never at one volume. A mixer rides it — down a good way under the first word of speech, held there for as long as the voice continues, brought back up in the pauses that are actually pauses, the ones between sections rather than between words. The listener never notices the moves, and that is the craft: the ramps are slow enough to be invisible, a fraction of a second of easing rather than a snap. The effect is exactly what the name suggests — the music ducks, the way a person standing between you and a stage crouches when the play begins. What the listener experiences is not “the music got quieter” but “the voice is important”. Done well, ducking is not a technical trick but a statement of priority, made continuously, sixty times a video, without a single conscious thought from the audience. Done by hand, it is also an afternoon of work per video, which is why automation exists.

Automation done crudely has its own failure mode, and it has a name: pumping. A naive ducker follows the voice too literally. Every gap in speech, however small, is read as silence, so the music surges up between your words and slams back down at the next syllable — up, down, up, down, in time with your breathing. The track stops being a bed and becomes a heartbeat, and the listener, who could not tell you what is wrong, knows that something is. The fix is a distinction the crude version misses: not all gaps are pauses. The space between two words is not a pause; neither is the beat where you take in air mid-sentence. A real pause — the end of a thought, a section boundary, the deliberate second of silence before a point — is long enough for the music’s return to mean something. The rule that separates good ducking from pumping is therefore simple to state: gaps shorter than a breath must stay ducked. Only when the silence is long enough to be intentional should the bed be allowed back up.

Talk2Camera’s editor builds that rule in. Lay a music track under the edit, switch on ducking, and the mix rides itself: while you speak, the music dips to about a quarter of the level you set for it, and in real pauses it comes back up, with ramps of roughly a third of a second so the moves ease instead of snapping. Gaps shorter than a breath stay ducked, so the bed never pumps along with your phrasing. What counts as “while you speak” comes from the take itself: on iPhone, from the take’s transcription; on Android, from analysis of the take’s own audio — which means it works even when you recorded without a script and there is no text to lean on. And like everything in the export pipeline, it is a decision rather than a commitment: the original recording is untouched, so you can export with ducking for the tutorial, without it for the trailer cut, or swap the music entirely, all from the same take.

The remaining decisions are yours, and they are musical rather than technical. Set the track’s level for the pauses, not for the speech — ducking will handle the speech; the level you choose is the level the viewer hears when you stop talking, and that is where a bed that is too loud gives itself away. Prefer music that can afford to be background: instrumental over vocal, since a second voice under yours is masking at its worst; steady over dramatic, since a track full of swells fights the ducker’s honest work with theatrical builds of its own. Then listen back once with your eyes closed. If you can follow every word without effort and still notice, in the pauses, that the video has a mood, the mix is right. If you notice the music while you are talking, it is too loud at its set level; if the pauses feel empty, it is too quiet. One pass of that test is usually enough — which, next to an afternoon of riding a fader by hand, is the entire argument.

Mentioned in this article

Music that ducks under your voice

Lay a music track under the edit and it dips while you speak and returns in the pauses — broadcast fader discipline, applied for you. Breath-length gaps stay ducked so it never pumps.

Attach an audio track

Lay music, a voiceover or a cleaned-up re-record under the edit, positioned where you want it. Preview and export both honour it.

Studio sound

One tap reduces room noise and evens out speech levels at export — a high-pass, a gentle gate and a slow leveller, tuned so nothing ever pumps. The recording itself is never touched.