Back to blog
5 min readCraft

Captions people actually read

Word-level timing was the easy part. The hard part is making captions feel like part of the video instead of a subtitle file bolted on top.

Open ten viral shorts and you’ll notice the captions before you notice anything else. Big, punchy, perfectly timed to the beat of the sentence — one or two words on screen at a time, never lagging, never racing ahead. That’s not a font choice. That’s a few hundred small decisions per minute of video.

Word-level timing is table stakes

Most caption tools time at the sentence or phrase level, then split the words evenly across that window. It looks fine until someone talks fast for one word and slow for the next — which is most sentences. Our subtitle agent reads word-level timestamps straight from the transcription agent, so each word appears exactly when it’s spoken, not when the average says it should.

The part nobody talks about: emphasis

Good caption styles don’t just display words — they highlight the ones that carry the sentence. “I built this in three days” reads differently than “I built this in three days” with no emphasis at all. The subtitle agent scores each word for stress using the audio’s pitch and loudness, not just its position in the sentence, and applies your brand’s emphasis style automatically.

The best caption style is the one you stop noticing — until you turn it off and the video suddenly feels slower.

Style is a memory problem

The other half of “captions people read” is consistency. If your last ten videos used a bold yellow keyword highlight and bottom-third placement, your eleventh video should too — without you re-selecting it. That’s why caption style lives in memory, not in a per-project settings panel. Set it once, and every new project starts from where the last one left off.

Captions in 48 languages, all with this same word-level timing and emphasis model — because the goal was never “captions exist.” It was “captions get watched.”