Talk2Camera
Menu
en-GB
video editing

How loud should speech be in a video?

Viewers forgive a dim picture and a wobbly frame. Speech they have to strain for, they punish within seconds.

Nobody rewatches a video because the picture was slightly dark. But dialogue that sits too low gets punished immediately: the viewer turns the volume up, gets a blast of room hiss along with your voice, then gets deafened by the next video in the feed. Most give up before trying. There is also no rescue coming from the platform. The big feeds apply loudness normalisation, but it mostly works downwards — a video that is too loud gets turned down, while a quiet one largely stays quiet. So the level you export is the level people hear, on a phone speaker, in a kitchen, over traffic. The working rule is blunt: speech should be the loudest thing in the mix by a clear margin, peaking high but never clipping, and it should stay that loud all the way through — not only in the sentences you delivered with confidence.

A music bed fails in one specific way: the moment a listener has to separate your words from the music, the music is too loud. It rarely sounds that way in the edit, because you know what the words are — your brain fills in whatever the mix buries. Two habits protect you. First, pull the bed down until it sounds almost too quiet on its own; under speech, that is usually about right, because voice and music fight over the same middle frequencies and the voice must win. Second, check the balance on the phone’s own speaker, not on headphones. A small speaker throws away the bass that made the bed feel gentle, and what remains sits exactly where your voice does. If the words survive there, they survive everywhere. And prefer instrumental beds — sung lyrics compete with speech directly.

All of this is easier when you can see the balance rather than hold it in your head. In Talk2Camera you can attach a track under a take — music, a voiceover, a cleaned-up re-record — positioned where you want it, and both the take’s sound and the attached track get a level line you drag across the waveform, the way a mixer works rather than a fader hidden in a menu. The bars redraw at the level you set, so a bed pulled down under speech looks quieter as well as sounding quieter: tall speech, low bed, visible at a glance. That picture is worth having, because a balance you can see is a balance you can check in two seconds before exporting, instead of listening to the whole video again with your eyes closed.

Two presenters bring one more problem: they are never equally loud. One sits closer to their microphone, one simply speaks more quietly, and a viewer riding the volume between two people gives up faster than a viewer riding it between videos. The fix has to happen before the voices are mixed together — once both people share one channel, no edit can pull one of them down without dragging the other along. Record through a dual wireless kit or a stereo interface and each presenter lands on their own channel; in the editor the take then splits into one lane per person, each with its own level line, so you can bring the louder voice down to meet the quieter one without touching it. The balance is set in the edit, where you can hear both — not lost at the receiver.

Mentioned in this article

Two-track audio with level lines

Your take’s sound and an attached track, each with a line you drag across its waveform — the way a mixer works, not a fader. The bars redraw at the level you set, so a bed pulled under speech looks quieter as well as sounding it.

Per-presenter levels

A two-presenter take splits into one lane per person, each with its own level line. Pull one presenter down without touching the other — the balance is set in the edit, not lost at the receiver.

Attach an audio track

Lay music, a voiceover or a cleaned-up re-record under the edit, positioned where you want it. Preview and export both honour it.