Talk2Camera
Menu
en-GB
content creation

The audience that reads: subtitles in a second language

A large share of your viewers never hear your voice. Another share understands it better in writing. One subtitle track serves both.

A talking video is consumed in silence far more often than the person who made it imagines. It plays on a train with the headphones left at home, in an open-plan office, in bed next to someone sleeping, in a feed that autoplays muted and never gets unmuted. The exact figures move around by platform and by year, but every survey that has looked lands in the same territory: a large share of feed viewing — in many settings a majority — happens with the sound off. Add the viewers who are deaf or hard of hearing, for whom captions are not a preference but the whole channel, and the conclusion is hard to escape. For a real and sizeable part of your audience, the captions are not an accessory to the video. They are the video’s audio, rendered in the only form that reaches them. A talking video without captions is, for these viewers, a silent film that forgot its intertitles.

Then there is the second audience: people who understand your language better in writing than in speech. Listening is the hardest language skill there is — it moves at the speaker’s speed, in the speaker’s accent, with the speaker’s swallowed syllables — and hundreds of millions of people have schoolbook reading fluency in a language they struggle to follow by ear. A caption removes the accent and half of the speed problem in one stroke. And past them stands the audience that does not read your language at all, for whom a translated track is the difference between watching and scrolling on. The economics here are lopsided in your favour. The video itself was expensive: the thinking, the script, the lighting, the takes. A second subtitle track costs almost nothing against that, and it widens the audience without re-recording a word — your voice, your face and your delivery stay exactly as they were.

The obvious rejoinder is that platforms already generate subtitles automatically, and they do — from the audio alone. The recogniser has never seen your script. It does not know that your product’s name has a capital letter in the middle, that your co-founder’s surname is not a common word, that your industry’s jargon is jargon rather than a mishearing. So it renders all of them phonetically, picks the wrong homophone with total confidence, sprinkles punctuation by statistics, and breaks lines wherever the model happens to breathe. When the platform then auto-translates those auto-captions, the errors compound: viewers in the second language receive a translation of a mistranscription, two steps from anything you said. And they do not blame the platform. A mistake in writing sits on screen, legible and attributable, next to your face. A slurred word passes in a second; a mangled name in a caption is a screenshot.

This is the case for shipping your own track, and it begins with where the text comes from. In Talk2Camera, caption text comes from your actual script or from on-device recognition of what you said — and when you read from the script you wrote, the captions inherit its spelling. Names come out the way you wrote them: your product’s odd capitalisation, your colleague’s surname, the technical term the autocaptioner has never met. That single property — the captions know what you meant to say — removes the whole class of error that makes automatic subtitles embarrassing. It also puts line breaks under your control rather than the model’s, which matters more than it sounds: a caption that breaks in the middle of a name, or strands a preposition, is read twice, and a caption read twice has already failed at its one job of being invisible.

For the second language, Talk2Camera translates your captions into another language as an additional subtitle track, and the detail that makes the feature work is timing. The translation keeps your original timing, line by line: the translated line appears exactly when you speak the original one. That sounds like a technicality until you watch a track that drifts. Your hands gesture at the thing you are naming two seconds after the reader learned about it, your face reacts to a joke the caption has not delivered yet, and the person on screen slowly detaches from the words below them. Line-by-line correspondence has a second, quieter benefit: it makes review possible. Anyone who reads the target language can lay the two tracks side by side and check them line against line — no timeline scrubbing, no guessing which sentence maps to which.

Shipping is a choice between two containers, and the app supports both. Captions can be burned into the picture, which is the right answer for feeds that ignore sidecar files: burned-in text is always visible, survives every player and every re-upload, and costs you the ability to turn it off or offer more than one language. Or they can travel as .srt files for the platforms that take sidecar subtitles, where they stay toggleable and searchable, and where a viewer can choose the language they want — the original and the translation can ride along together. The practical rule: burn for short vertical video, ship .srt where the platform respects it, and do both when a video lives in more than one place. However it travels, the point stands. The viewer who reads is not an edge case. On the evidence, they may be your majority — and they deserve better than a phonetic guess at your own name.

Mentioned in this article

Caption translation

Turn your caption track into a second .srt in another language — same timings, one line per line, never merged or split. Part of a subscription’s AI allowance.

Export subtitles as .srt

A separate subtitle file for platforms that want one, instead of burning the words into the picture.