A talking-head video is almost never just talking. Somewhere in the middle you have to hold up the product, point at the line on the chart that proves your point, walk to the window, put on the glasses you have spent the last minute describing. These moments are usually the reason the video exists at all — nobody plans a demonstration in order to avoid demonstrating — and yet they are the one part no script quite knows how to carry. Words go in the script. Actions go somewhere else, and somewhere else almost always means your memory. So presenters end up running two systems at once: a prompter that reliably holds every sentence, and a head that is supposed to hold every move. One of those systems never forgets anything. The other is the one that fails at minute three, when you sail straight past the moment the product should have come up and only notice while reviewing the take, box still sealed on the desk beside you.
Memory fails here for a boring, mechanical reason, not a personal one. Delivering a script takes nearly everything you have: reading ahead, sounding unrehearsed, keeping the gaze somewhere near the lens, managing breath. The plan to pick up the box at the right sentence is stored in the same head that is doing all of that work, and it surfaces exactly when the load lightens — before the take begins, and again in the second after the take ends. During, it is simply gone. This is why a forgotten action feels so absurd afterwards: you knew it cold at breakfast and you know it cold now, holding the box you never opened. Rehearsal narrows the gap but never closes it, because the thing that crowds the plan out of your head is not unfamiliarity with the script. It is the performance itself, and the performance is present in every take you will ever record, including the good ones.
The obvious fix is to write the actions into the script, and everyone who tries it discovers the same trap: you read them aloud. This is not carelessness. A prompter trains exactly one reflex — words arrive, you say them — and that reflex does not stop to ask whether a line was meant for the audience or only for you. At reading speed there is no spare attention for evaluating each sentence as it arrives; not having to evaluate is the entire service a prompter provides. Capital letters help less than you would hope, and so do brackets, because at a glance they are still text sitting in the column of text you have agreed to speak. On a good day you catch yourself two words in and grimace. On a bad day the take ends with you instructing four thousand viewers, in your warmest presenting voice, to show the product now.
What the script needs is a second kind of line — one the prompter itself understands is not speech. In Talk2Camera, a line written like [SHOW THE PRODUCT] becomes a stage-direction chip in the prompter: it scrolls up with the script, sitting visibly apart from the sentences around it, and its entire job is to be shown, acted on, and never spoken. The reading reflex gets nothing it recognizes as prose, so the mouth leaves the chip alone while the hands do what it says. You write the direction exactly where it belongs — between the sentence you say before the action and the sentence you say after it — which is precisely the place your memory kept losing it. The choreography and the words live in one document, in order, each move filed at the exact sentence it serves. And of everything present on a set, the document is the one participant that never gets nervous.
Where a stage direction never goes matters as much as where it appears. A chip cannot reach voice follow: the speech recognizer never waits for words you were never going to say, so the scroll keeps tracking your voice through the action instead of stalling on a line nobody will read out. And a chip cannot reach the captions — the caption track stays clean, which anyone who has spotted a bracketed instruction burned into a finished video will recognize as the difference between a note to self and a published mistake. That separation is the whole idea. A script that mixes saying with doing has to keep the two apart everywhere the text flows: on the prompter, in the scroll, in the subtitles. A direction that stays visible to exactly one audience — you — is the only kind that is safe to write down at all, and safety is what lets you write down as many as the shoot really needs.
A few honest limits. A chip reminds you to do the thing; it does not rehearse the thing for you, and an awkward manoeuvre — unboxing something taped shut, plugging in a cable on camera — still deserves one dry run before you roll. Directions work best short and imperative, SHOW THE PRODUCT rather than a paragraph about how to show it, because you will be reading them at a glance in mid-performance. And a script with its choreography worked out is worth keeping: in Talk2Camera, scripts live in a library with favourites and folders, and each one carries a word count and an estimated read time, so the version with the directions in place is the version you find again next month, not a draft stranded on another device. The move you almost forgot this time is written down for every take after it — which is the quiet advantage of treating actions as part of the script instead of a promise you make to yourself.