Back to blog
How-To

Caption Timing and Reading Speed Guide

Bipul Kumar

Caption Timing and Reading Speed Guide

Good captions are synchronized, readable, and faithful to the speaker. A technically valid SRT file can still be exhausting to follow when cues overlap, flash too quickly, or contain long unbroken lines.

This guide is a review pass you can run before export: check timing, measure reading speed, keep lines scannable, then preview against the audio.

Caption review workflow showing a timeline, stacked subtitle blocks, a stopwatch, and simplified lines
Time the cues first, then check whether a viewer can finish each line before it disappears.

Check the timing first

Every cue needs a start time before its end time. Adjacent cues should not overlap unless the file deliberately represents two speakers. As a useful starting range, keep most cues on screen for one to seven seconds. Listen around every boundary rather than judging timestamps only from a text list.

Timing errors to catch immediately:

  • A cue that starts after it ends.
  • Two cues that occupy the same moment for one speaker.
  • A cue that appears a beat after the spoken word, so the viewer reads late.
  • A cue that lingers into the next sentence.
  • A one-frame flash that nobody can read.

If you generated captions from a script, the words may be correct while the timestamps are not. Treat timing as its own pass. Do not “fix wording” until the file is in the right place on the timeline.

Measure reading speed

Characters per second gives a practical warning when too much text is squeezed into a short cue. Start with 20 characters per second for Latin-script captions and 12 for Chinese or Japanese. These are review thresholds, not universal laws. Simplify the wording or extend the cue when viewers cannot finish comfortably.

A fast way to review:

  1. Note the on-screen duration of a dense cue.
  2. Count characters in that cue, including spaces.
  3. Divide characters by seconds.
  4. Flag anything above your threshold.
  5. Either cut words or give the cue more time.

Do not solve a reading-speed problem by shrinking the font in a burned-in subtitle. The caption file should remain readable at a default player size.

If the speaker is fast, you still have three honest options: paraphrase, split across two cues if the audio allows, or accept that some asides will be shortened. Captions that omit a joke’s wording but keep its meaning are often more accessible than a wall of text.

Keep lines scannable

Start with a maximum of 42 characters per line for Latin scripts and 20 for Chinese or Japanese. Break lines at natural phrases. Do not split a name, number, or tightly connected phrase merely to reach an exact width.

Prefer two short lines over one long line. Keep the second line shorter or equal when you can. Avoid ending a line on a weak word such as “of,” “the,” or “and” if the next line holds the meaning.

Speaker labels, sound descriptions, and italics for emphasis should stay consistent. If you mark `[door closes]`, mark similar events the same way throughout. Inconsistent conventions look like errors.

Fix safely

Automatic cleanup can repair numbering, whitespace, and timestamp separators. Timing and wording require a person. Preview the corrected captions against the audio, export a new SRT or VTT, and preserve the source file in case the edit needs to be reversed.

A safe edit sequence:

  1. Duplicate the caption file before any rewrite.
  2. Fix invalid timestamps and overlaps.
  3. Repair reading speed and line length.
  4. Check names, numbers, and on-screen text.
  5. Play the video with captions on, at least around every edit.
  6. Export a new file with a dated name.

Use the SRT to speech workspace when you also need to audition caption text as narration. Hearing the captions out loud is a fast way to catch a line that is technically timed but impossible to follow.

If you are producing a YouTube voiceover, generate captions from the approved script rather than from a fuzzy transcription of the final mix. The YouTube voiceover workspace can export cues alongside the audio so timing starts from the same source.

What “good enough” looks like

Accessible captions are not a transcript dump. They are a reading experience timed to speech.

  • A viewer can finish each cue without rushing.
  • Two speakers are distinguishable.
  • Sound that matters is described briefly.
  • Names and numbers match the programme.
  • The file plays in common players without overlapping itself.

If you only have time for one pass, do timing plus reading speed. Wording polish can wait; unreadable timestamps cannot.

Frequently Asked Questions

### What reading speed should I use for English captions?

Start around 20 characters per second as a warning threshold. If a cue exceeds that, shorten the wording or extend the duration. Children’s content, language learning, and complex technical speech often need more time.

Should captions match the speaker word for word?

Match meaning first. Verbatim captions are valuable for legal, educational, and quote-heavy material. For entertainment, a slight paraphrase that preserves names, numbers, and intent is often more readable.

Is SRT or VTT better?

SRT is widely supported. VTT is better when you need web-native styling or additional metadata. Export the format your platform accepts, and keep a clean SRT as a fallback.

Can I auto-fix caption timing?

You can auto-repair numbering and obvious timestamp syntax. Overlaps, reading speed, and line breaks still need a person listening to the programme.

Try it yourself

Convert text to speech free. No signup, no fees.

Open the Converter
visual guide showing PDF, DOCX, EPUB, TXT, HTML, Markdown, and subtitle files converting into audio

Visual guide

A document-to-audio workflow for listening to files, articles, books, and notes.