Back to blog
Accessibility

An Accessible Audio Description Workflow

Bipul Kumar

An Accessible Audio Description Workflow

Audio description communicates important visual information to people who are blind or have low vision. A useful track explains what matters without narrating every detail or competing with dialogue.

This is a production workflow, not a slogan: map the gaps, write observable description, time the language, generate a distinct stem, then review it against the programme.

Audio description workflow with film frames, timed gaps, and a narration stem mixed under dialogue
Write only into the gaps. Time each cue, generate a separate stem, and review it against dialogue and essential sound.

Map the available spaces

Watch the programme and mark gaps between dialogue, music cues, and essential sound effects. Record the maximum duration available for each description. Timing determines how much language the script can contain.

A gap is not “any quiet.” A held look, a door slam that tells the story, or a musical sting may need to stay uncovered. If covering that sound would hide the plot, the description belongs somewhere else or must be shorter.

Build a cue sheet:

  • Timecode in
  • Maximum seconds
  • What must be described
  • What must not be covered
  • Draft sentence

If there is no gap, you cannot invent one with a faster voice without harming the soundtrack. Redesign the sentence or wait for the next opening.

Describe observable information

Prioritize actions, expressions, locations, on-screen text, and visual changes needed to understand the story. Use concise present-tense language. Avoid interpreting a character’s intention when the image does not establish it.

Prefer “she folds the letter and puts it in the drawer” over “she sadly hides the letter because she cannot cope.” The second sentence is a guess. The first is what a viewer with sight would be able to see.

On-screen text is often essential: caller ID, a news headline, a password on a prop laptop, a street sign. If the plot depends on reading it, describe it. If it is texture, you can skip it when time is short.

Do not describe every costume change. Do describe a costume change that identifies a disguise.

Test duration and pronunciation

Use the speech time calculator for each cue. Preview character names and specialist terms before generating the full track. Choose a clear voice that remains distinct from the programme dialogue without drawing attention away from it.

If the calculator says the sentence is 4.2 seconds and the gap is 3.0, cut words. Do not crank speed until the description sounds like a disclaimer. A slightly incomplete description that fits is more usable than a perfect sentence that talks over a confession.

Fix names with the same method as any other TTS job: phonetic cues, then a one-sentence preview.

Generate description as its own stem, not mixed into the show audio. You need to duck, move, or rewrite one cue without re-exporting the whole programme.

Review with the programme

Place each generated segment against the original audio. Check that descriptions do not cover speech or meaningful sounds. Invite blind or low-vision reviewers whenever possible; technical timing alone cannot prove that the information is useful.

Review questions:

  • Can a listener follow the plot without the picture?
  • Did we talk over dialogue?
  • Did we hide a sound that is acting as a story beat?
  • Are names consistent with the main track?
  • Is the description voice obviously separate, but not cartoonish?

Export separate narration stems and a timing manifest so an editor can mix, revise, and audit every description independently. Keep the cue sheet with the stems. The next revision should not start from a bounced mix.

Where TTS helps, and where it does not

Text to speech is useful when you need consistent delivery, fast iteration, and a stem you can regenerate after a wording change. It is not a substitute for the editorial judgment of what to describe.

If the programme already has a human description track you are allowed to use, prefer that unless you are producing a new, rights-cleared version. If you are describing your own video, TTS can get a first pass in place so reviewers have something to react to.

Disclose synthetic description when listeners might assume a person recorded it, especially in education, broadcasting, and public services.

Frequently Asked Questions

### How much should I describe?

Enough for the story and the information a sighted viewer would use. Not every visual detail. Time in the soundtrack is the hard limit.

Can I describe over music?

Sometimes, if the music is background and the gap is real. Do not cover lyrics or a motif that is acting as narrative. When unsure, leave the music and shorten the sentence.

Should audio description match caption wording?

They solve different jobs. Captions represent speech and relevant sound. Description represents essential visuals. They should not contradict names, places, or on-screen text.

What should I deliver to an editor?

A dry description stem, a cue sheet with timecodes, and notes on gaps you left intentionally. A mixed file hides the decisions.

Try it yourself

Convert text to speech free. No signup, no fees.

Open the Converter
visual guide showing text to speech resources, voice testing, support, and helpful guide content

Visual guide

A knowledge guide for text-to-speech support, tutorials, and editorial resources.