Video captions: timing, formats and accessibility
A transcript is read as a document; captions are read while watching. Accurate words still need short, readable and synchronised cues.
Measure caption quality
A cue should appear with speech, remain readable and avoid covering important visuals. Long timestamped paragraphs are poor captions even when every word is correct.
Line length, cue duration, reading speed and shot changes affect comprehension independently of whether the recognised words are correct.
Choose the delivery format
SRT is widely supported, while VTT is designed for web media and richer cues. Separate caption tracks remain switchable and accessible; burned-in text does not.
Keep an editable caption file as the source and produce burned-in versions only as distribution derivatives for channels that require them.
Review automatic output
YouTube itself recommends reviewing automatic captions because accents, noise and overlapping speech cause errors. Check names, line breaks and cue timing with the video.
Review specialist vocabulary, speaker overlap and music transitions against the final video, not only against an isolated audio export.
Pair captions with a transcript
W3C treats captions and transcripts as related but distinct accessibility components. A linked transcript also supports search, quotation and direct navigation.
Publish a nearby transcript when appropriate so viewers can search and quote the material beyond the constraints of timed cues.
Continue with the underlying material
Need an estimate for your project?
Tell us about the project. We will break it into stages and explain the budget drivers.