Every few years an editing idea arrives with a pitch attached, and the pitch for text-based editing is that the timeline is finished — that scrubbing through footage will soon look the way rewinding a tape looks now. Timeline editors return the favour by treating text editing as a toy for people who never learned the real tool. Both camps are arguing about the wrong thing. The two views are not rival answers to one question; they are answers to different questions, and an ordinary talking-head edit asks both kinds before it is done. The useful skill is not loyalty to either view. It is knowing which question you are answering at any given moment — because a meaning question answered on a timeline wastes an afternoon, and a frame question answered in text cannot be answered at all. Sort your edit into those two piles and most of the argument dissolves. What is left is a workflow, and it is a calm one: each view doing the work the other is bad at, in an order that turns out to matter.
Start with what the timeline is genuinely for, because the list is longer than the text-editing pitch admits. Where does the video begin? The recording starts with you reaching for the button and settling your face; the first word comes a second and a half in, and choosing the exact frame to open on is a decision about frames, not about words. A transcript has no vocabulary for it — it cannot show you a blink, a hand leaving the shot, the half-second where your smile sets. The same goes for tightening without removing: two good sentences separated by a breath that is a beat too long, where the fix is to shorten a gap by ten frames, not to delete anything a transcript could name. And it goes for sound: a seam that lands cleanly on a consonant, a breath faded rather than chopped, a hum you can see in the waveform and never in the words. These are questions about time and sound. The timeline is not a legacy interface for them; it is the correct instrument, and it will stay the correct instrument.
Now the questions text is for. A sentence that is simply wrong — the flubbed first attempt, the number you corrected a moment later — is a meaning problem, and hunting for it by scrubbing is archaeology: you dig through strata of grey waveform for something you could recognise instantly by sight. The tangent that seemed worth taking and was not; the aside that blunts the strong point right before it lands; the paragraph where you said the same thing twice and the second version is better — these are editing decisions of the kind an editor of prose makes, and they are best made the way prose is edited: by reading. In text, each of them costs one read and one tap. And the deeper effect is not speed but coverage. At reading speed you weigh every sentence in the take, the way you weigh every sentence of a draft; at scrubbing speed you only fix what was catastrophic enough to remember. Text editing does not just find the cuts faster. It finds cuts the timeline would never have surfaced at all.
In Talk2Camera the choice is not exclusive, because the two views edit one shared cut list. The Transcript tab lists the take as tappable sentences — on iPhone and iPad transcribed on the device, on Android taken from the script the take was read from, timed by the prompter as you spoke. Tap a sentence, cut it, and that stretch of video leaves the edit; the sentence stays visible, struck through, and restores with a tap. Switch to the timeline and the same cut is simply there, a gap you can inspect and play across — and the trims you make on the timeline, the frame-level decisions about where things begin and end, belong to the same single edit the transcript is showing you. Every change, from a cut sentence to a trimmed edge to the pauses removed automatically, lands in one reviewable list where each item can be undone on its own. And beneath both views, the recording file itself is never modified. There is no export from one mode to the other, no point of no return — just two windows onto the same set of decisions.
That shape suggests an order of work, and it is worth following. Do the text pass first. Read the take, cut the sentences that should not exist, and do it without worrying about seams — there is no sense in frame-polishing a sentence you are about to delete. Then walk the timeline over what survived: set the opening frame, tighten the breath between the two sentences that now sit together, check the seams the text pass created and nudge the one that clicks. The first pass is editing as writing; the second is editing as craft. Ten minutes of each will beat an hour of either alone, because each pass only faces the questions it is equipped for. The camps will go on arguing, since a pitch needs an enemy. But an edit of a person talking is both a piece of writing and a piece of footage, and it seems unremarkable, once you have worked this way, that it should be edited as both.