Talk2Camera
Menu
en-US
video editing

Music under speech: choosing a bed that behaves

The best music in a talking video is the track nobody can hum afterwards. And then someone still has to ride the fader.

Music under speech has a job title in broadcasting — a bed — and the name is the brief. A bed is what the voice lies on. It is not the thing you look at, and the moment it becomes the thing you listen to, it has failed at the only task it had. In a talking video the bed does three quiet jobs. It gives the piece a floor of energy, so that a pause in the speech does not open onto silence and make the room suddenly audible. It carries continuity across cuts, so that two takes shot an hour apart feel like one sitting. And it sets a temperature — the same sentence reads as confident, wistful or urgent depending on what is underneath it. The test for whether a bed did its work is unglamorous: ask a viewer afterwards to hum it. If they can, the music was competing; if they remember that there was music but cannot reproduce a bar of it, the bed behaved. You are not scoring a film. You are upholstering a voice.

What survives under speech is almost never what sounds best on its own. The tracks that work are sparse — few instruments, space between the notes, arrangements that repeat without developing much, because development pulls attention and attention is the one budget the voice cannot share. Vocals are the clearest exclusion: two voices speaking different words put the listener’s language system in a tug of war, and even distant background vocals or a chopped vocal sample will make a viewer feel they missed something. Tempo matters in a subtler way. A bed noticeably faster than your natural speaking rhythm makes you sound like you are lagging behind your own video, and a crawling one drags the energy of a brisk explainer down with it. You do not need to match beats per minute to words per minute; you need the music to feel like it is walking at your pace. Instrumental, repetitive, patient — the qualities that would make a track boring at a party are precisely the qualities that make it employable under a voice.

Where beds go wrong is rarely loudness alone; it is real estate. A speaking voice lives in a narrow band — its weight sits low, roughly where a cello sits, and its intelligibility, the consonants that let a listener tell one word from another, lives in the presence region a few octaves up. A track whose activity is concentrated in that same middle territory fights the voice for the same square metres, and no fader setting fully fixes it: turn it down enough to stop the masking and it disappears, turn it up enough to hear and the words blur. This is why strummed acoustic guitar, dense synth pads and busy piano — the staples of every royalty-free library — are the worst offenders, and why they audition so innocently: soloed, they sound warm precisely because they are rich in the middle. What coexists with a voice is what lives above and below it — a bass line, a kick, hi-hats, airy chimes, sparse plucks high up. Choose by listening to the midrange with your own speech playing, not by how the track feels alone.

Then comes the discipline, which is mostly the discipline of restraint. Set the bed lower than feels right — meaningfully lower, because your ear, proud of having chosen the track, wants to hear it, and the viewer has no such loyalty. Check the mix on the worst speaker you own, which is the phone in your hand: small drivers exaggerate exactly the midrange collision you are trying to avoid. Keep the level consistent across the whole piece rather than nudging it scene by scene, because a bed that wanders reads as a mistake even when each level was chosen deliberately. And notice what happens at the edges of speech. When you talk, the music must be far down; when you genuinely stop — a section break, a held look — it should come back up and briefly carry the piece. In broadcast, that movement was a person’s job: an operator sat with a hand on the fader, pulling the music down as the presenter breathed in and easing it back in the gaps. Which raises the practical question for anyone working alone: who rides the fader?

In Talk2Camera, the answer is that the ride is automated and the choice stays yours. You attach the track and set its level — one line, one decision. With ducking on, the bed drops to about a quarter of that level whenever you speak, and comes back up in real pauses, with ramps of about a third of a second, so the movement is a breath rather than a jolt. The breath-length gaps inside sentences stay ducked — the bed does not surge up between every phrase and dive down again, the pumping that gives automated mixes away. Only pauses long enough to be structural let the music rise. All of it happens at export, like everything else in the app’s finishing pass: the recording itself is untouched, so you can re-export with a different track, a different level, or no bed at all, and the voice — cleaned by studio sound on the way through — is the one element that never has to be performed again. The bed is a costume. The recording underneath stays yours.

Mentioned in this article

Music that ducks under your voice

Lay a music track under the edit and it dips while you speak and returns in the pauses — broadcast fader discipline, applied for you. Breath-length gaps stay ducked so it never pumps.

Attach an audio track

Lay music, a voiceover or a cleaned-up re-record under the edit, positioned where you want it. Preview and export both honour it.

Studio sound

One tap reduces room noise and evens out speech levels at export — a high-pass, a gentle gate and a slow leveller, tuned so nothing ever pumps. The recording itself is never touched.