Two-host scripts usually fail through politeness. Each host is given a paragraph, delivers it whole, and hands the floor to the other, who has been nodding through it and now does the same. That is not a conversation; it is two monologues taking turns, and viewers hear it in seconds, because real conversation is not polite. Real turns are short — a sentence or two. People finish each other’s setups, answer questions with questions, and react in the middle of the other person’s point. So write the exchange as an exchange. Give the hosts different jobs: one carries the argument, the other carries the audience, asking what a viewer would ask and doubting where a viewer would doubt. Keep most turns under twenty words. Give reactions their own lines instead of trusting them to happen. Then read it aloud together once before the shoot — every line that trips one of you was written for the page, and you will hear which.
The format that carries all this is old, plain, and not worth improving: start each line with the speaker’s name and a colon. Maya: writes Maya’s lines. Sam: writes Sam’s. It survives every copy and paste, needs no special software, and has the one property that matters at a glance — the hand-overs are visible down the left margin. That margin is also your balance check. If one name owns the column, you have written a lecture with a witness, and it is cheaper to discover that in the draft than in the footage. Keep each turn on its own line rather than folding two voices into a paragraph; the format is doing your blocking for you.
On the prompter, the same tags become behaviour. Talk2Camera reads Name:-tagged lines as a dialogue script: each speaker’s lines take their own colour, and every hand-over is announced, so neither of you counts lines ahead to find out when you are next. Voice follow stays exact throughout, because the name tags are never fed to the speech recogniser — the script tracks the words actually being said, whichever of you is saying them. If you draft with the AI writer, a toggle has it write in this format from the start. The practical effect is bigger than it sounds: both hosts stop watching the script for their cue and start listening to each other, and listening is where the conversational sound actually comes from.
Record the two voices separately, or regret it in the edit. Two hosts on one track cannot be balanced afterwards — raise one and the other rises with the room. A dual wireless kit solves this at the point of capture: a RØDE Wireless GO II, or a DJI Mic switched into split mode, puts each transmitter on its own side of a stereo pair instead of mixing them into one. Plug the receiver into the phone, and in Talk2Camera selecting it as the input is the whole setup — each presenter lands on their own channel. Any stereo interface behaves the same way. The built-in microphone, deliberately, stays mono: one phone mic cannot honestly separate two people, and a pretend split would be worse than none.
The payoff comes in the edit. A two-presenter take splits into one lane per person, each with its own level line, so the quiet host comes up and the loud one comes down without either dragging the other along — the balance is set in the edit rather than lost at the receiver. It also forgives the classic dialogue accident, both of you laughing over the punchline: on one track that moment is fixed forever, on two you choose whose laugh survives. A conversation is worth scripting. It is also worth recording as two people, because that is what it was.