Talk2Camera
Menu
en-GB
video editing

Cut, dissolve or dip to black: how to join takes

Every join between two takes makes a claim about time. A hard cut, a cross dissolve and a dip through black each make a different one.

Record a video in two takes, butt them together, and something small but wrong happens at the seam: you teleport. Your head is suddenly two degrees to the left, your hands have jumped from the desk to your lap, and the viewer’s eye — which was not consciously watching either — registers a glitch in the world. This is not pickiness on the audience’s part. The visual system is built to track continuity, and a century of film grammar has trained everyone to read a join between two shots as a statement about time. A cut between two different framings says “same moment, new angle”. The same cut between two nearly identical framings says “time is missing here”, and the body’s little teleport is the evidence. None of this makes joins bad. Every edited video is made of them. It only means a join is never neutral: whichever way you do it, you are telling the viewer something about what just happened to time, and the only real mistake is telling them something you did not mean.

The hard cut is the default, and its meaning is the bluntest: next instant, keep up. Between two genuinely different shots it is invisible — the change of framing explains the change of image, and the viewer reads continuity straight through it. Within one framing, the hard cut became the signature of the vlog: the jump cut, once an error, now a style that says “I removed the boring part and we both know it”. That style works on rhythm and honesty. A video cut that way jumps every few sentences, consistently, from its first minute — so the viewer is taught the convention early and stops seeing it. Where the hard cut fails is in the middle ground: a video that flows smoothly for three minutes and then jumps once, visibly, reads not as style but as a mistake, one splice sitting proud of an otherwise seamless surface. The rule of thumb is symmetry. Either cut hard often enough that jumping is plainly the video’s language, or make the one join you truly need with something softer than a cut.

The two softer options say different things, and they are not interchangeable. A cross dissolve overlaps the outgoing and incoming picture, and it reads as continuity of subject across a gap in time: same person, same thread, a little later. It is the join for two halves of one argument recorded on different days, for stitching a re-recorded section into an old take, for any seam you want the viewer to glide over rather than notice. Its danger is sentimentality — a dissolve every thirty seconds turns a talking head into a slideshow. A dip to black is a different tool entirely, and stronger. The picture fades all the way out, the screen is genuinely empty for an instant, and the next section fades in. That empty instant is punctuation: not a comma but a full stop, the visual equivalent of a chapter break. It tells the viewer the previous thing is finished — we are done with the introduction, done with the theory, and what comes next is a new room. Give them that breath at real boundaries and the video feels structured. Give it to them mid-thought and it feels like a power failure.

Talk2Camera treats the join as the one decision it is. Merging several recordings into one video, you choose per seam: hard cut, cross dissolve, or dip to black. The dip is honest about its punctuation — it passes through a true black frame at its midpoint, so the screen actually empties rather than merely dimming. The length of a transition is adjustable, and it is bounded by a rule that protects short material: a transition never eats more than forty percent of the shorter neighbouring take, so a four-second pickup cannot be swallowed whole by the dissolves on either side of it. And underneath every join, whatever the picture is doing, the audio crossfades — the sound of one take eases into the sound of the next instead of clicking across the splice, because rooms never sound quite identical twice, and an audible seam betrays a join faster than a visible one.

In practice, talking-head video needs each of the three, in different places. The flubbed line you re-recorded as a pickup wants a hard cut at a sentence boundary — the join hides in the pause, and softness would only draw attention to it. The two halves of a tutorial recorded on different mornings want a short cross dissolve, so the light shift and the slightly different posture glide past instead of jolting. The move from your introduction to the screen demo, or between numbered sections of an argument, wants the dip to black — it is the moment the viewer is allowed to reset, and the structure of the whole video gets clearer every time you grant it honestly. Then be consistent. Pick the video’s grammar in the first minute and keep it; a piece that cuts hard, dissolves and dips at random is mumbling in three languages. And when you genuinely cannot decide, cut. The transition is only punctuation, and punctuation has never rescued a sentence. The words carry the video; the joins just have to stay out of their way.

Mentioned in this article

Merge takes

Join several recordings with a hard cut, a cross dissolve, or a dip through black — a true black frame at the midpoint, the chapter break of talking-head grammar. Audio crossfades under the join.

Trim on a zoomable timeline

Pinch to zoom. At full magnification one pixel is milliseconds, so an in-point lands on the frame you meant.