Captions, transcripts, and audio description are related accessibility tools, but they are not interchangeable. Each preserves a different part of an audiovisual experience and supports a different mode of access.
A useful media-accessibility process begins by asking two questions: What information exists in the audio, and what information exists only in the image? The appropriate alternatives should preserve both the content and the task, not merely satisfy a generic requirement to attach text to a video.
Begin with a content inventory
Consider a short instructional video in which a presenter explains how to replace an air filter. The narration describes opening the access panel and removing the old filter. An arrow on screen shows the correct airflow direction, a model number appears briefly, and an alarm sounds when the panel is not fully closed.
Different alternatives preserve different parts of that experience:
- Captions synchronize the presenter’s speech and the meaningful alarm sound with the video.
- A basic transcript makes the speech and relevant sound information available as readable text.
- A descriptive transcript can also include the airflow arrow, model number, and other important visual information.
- Audio description communicates the visual-only instructions at the appropriate moments during playback.
Captions alone would not communicate the direction of the on-screen arrow unless that information were also spoken. A transcript containing only dialogue would omit the same instruction. The first production task is therefore not choosing a file format. It is identifying where meaning lives.
Captions preserve synchronized audio information
Captions present speech and meaningful non-speech audio as text synchronized with the media. They support people who are Deaf or hard of hearing, and they can also help people watching in noisy or quiet environments, readers processing unfamiliar terminology, and anyone reviewing dense material.
Depending on the content, effective captions may include:
- spoken dialogue and narration;
- speaker identification when the speaker is not visually apparent;
- alarms, knocks, machinery sounds, or other sounds needed to understand events;
- music or lyrics when they carry meaning;
- changes in tone that are necessary to interpret the statement; and
- appropriate punctuation, timing, and line breaks.
Caption timing matters. Text should appear closely enough to the corresponding audio that a viewer can follow both without having to reconstruct the sequence. Captions should also remain on screen long enough to be read without obscuring essential visual content.
Closed captions and open captions
Closed captions can be turned on or off through the media player. They are commonly delivered as a separate timed-text track. Open captions are embedded directly into the video image and remain visible to everyone.
Open captions can be useful when playback systems do not reliably support caption tracks, but they give viewers less control. They may also become difficult to read when a video is reduced in size. Closed captions generally offer more flexibility when the player exposes them accessibly and renders them reliably.
Captions and subtitles are not always equivalent
Captions and translated subtitles may use similar technical systems, but they do not necessarily contain the same information. Subtitles commonly translate or reproduce speech for viewers who can hear the soundtrack. Accessibility captions also communicate meaningful non-speech audio and identify speakers when necessary.
A translated subtitle track should not be assumed to satisfy the need for captions if alarms, music, speaker changes, or other important audio information have been omitted.
Automatic captions require human review
Speech-recognition systems can provide a useful draft, particularly when audio is clear. They can also misinterpret names, accents, technical terminology, model numbers, punctuation, and overlapping speakers. They generally cannot determine consistently which background sounds carry meaning.
Review automatic captions against the final media. Correcting visible words without reviewing timing, speaker changes, and meaningful sounds leaves important accessibility work unfinished.
Transcripts provide a readable representation of the media
A transcript presents media content as text that can be read independently of synchronized playback. It supports people who cannot access the audio, but its usefulness extends further. Transcripts make media easier to search, quote, translate, study, review, and reference on a slow connection or in an environment where playback is inconvenient.
Basic transcripts
A basic transcript usually includes speech, speaker identification, and meaningful audio information. It is well suited to audio-only content such as interviews, podcasts, recorded meetings, and lectures when it preserves all information needed to understand the recording.
Descriptive transcripts
A descriptive transcript also includes important visual information. This may include:
- actions and demonstrations;
- expressions or gestures that change meaning;
- on-screen text;
- charts, diagrams, and visual comparisons;
- scene changes and locations;
- the identity of a person who is visible but not named in the audio; and
- visual instructions required to complete a task.
A descriptive transcript can provide a complete text-based representation of a video, but it does not automatically replace synchronized captions. Someone watching a video while relying on text for the soundtrack needs captions at the time the corresponding events occur.
Publish transcripts where people can find them
Whenever practical, provide the transcript as accessible HTML near the media. HTML can adapt to viewport size, respond to user text settings, support heading navigation, and remain available to browser search and assistive technology.
A downloadable document may be offered as an additional format, but it should not be the only option without a clear reason. Label transcript links plainly, identify the format when relevant, and avoid burying the transcript behind ambiguous text such as “more” or “download.”
Audio description communicates visual-only meaning
Audio description is spoken narration of important visual information that the existing soundtrack does not convey. It can support people who are blind or have low vision by making actions, expressions, settings, diagrams, on-screen text, and other visual meaning available through audio.
Useful audio description may communicate:
- who enters or leaves a scene;
- an action that is not explained in dialogue;
- a facial expression that changes the meaning of a response;
- text, labels, or instructions appearing only on screen;
- the relevant movement in a demonstration;
- changes shown in a chart or diagram; and
- the identity of a speaker who is visually apparent but not named.
The goal is not to narrate every visible detail. Description should communicate the visual information necessary to understand the content, context, or task. Excessive description can compete with dialogue and make the media harder to follow.
Standard and extended audio description
Standard audio description is normally inserted into natural pauses in the existing soundtrack. If the media contains too few pauses to describe essential visual information, an extended described version may pause the original program long enough to provide the needed explanation.
Description can be delivered through a separate audio track, a separately published described version, or narration integrated into the original production. Delivery support varies among browsers, platforms, and players, so the completed experience should be tested rather than inferred from the presence of a file.
Accessible production can reduce avoidable visual gaps
Presenters can often express important visual information naturally while recording. Instead of saying, “Move this over here,” a presenter might say, “Move the airflow switch to the left-hand position marked ‘intake.’” This improves the original narration for many listeners and may reduce the amount of additional description required.
This approach should not become unnatural over-narration. It is a planning practice: communicate essential information through more than one sensory channel when doing so fits the material.
Choose alternatives according to the media
The necessary combination depends on whether information is carried through audio, video, or both.
Prerecorded video with speech or meaningful sound
Provide synchronized captions. If important visual information is absent from the soundtrack, provide audio description or another appropriate media alternative according to the applicable requirements. A descriptive transcript is also valuable for reference and text-based access.
Audio-only content
Provide a transcript containing speech, speaker identification, and meaningful sounds. A transcript is usually more appropriate than captions because no visual timeline needs to be followed, although a platform may display synchronized text as an additional feature.
Video-only content
Provide an alternative that communicates the important visual information. Depending on the content, this may be an audio track, descriptive transcript, or concise text alternative. A silent decorative animation may not need a detailed equivalent, but it should not create distraction or interfere with access.
Live audiovisual media
Live captions require trained captioners, real-time speech recognition, or a combination of technologies and human oversight. Production conditions differ from prerecorded media, and small delays may be unavoidable.
If a live event is later published as a permanent recording, review and correct the captions. An unedited live-caption stream may contain errors, omissions, timing problems, or event-specific text that does not work as an archival alternative.
Media that demonstrates a task
An alternative should preserve enough information for someone to complete or understand the task—not merely summarize the topic. Include measurements, sequence, warnings, control names, visual states, and other details that affect the outcome.
Accessible files require an accessible media player
A correct caption or description file does not help if the viewer cannot find or operate it. The player is part of the accessible media experience.
Media controls should provide:
- full keyboard operation;
- visible keyboard focus;
- accessible names for play, pause, volume, captions, descriptions, playback speed, and full-screen controls;
- clear state information, such as whether captions are on or off;
- a logical focus order;
- sufficiently large control targets;
- compatibility with screen readers and other assistive technologies;
- caption text that remains readable at different viewport sizes; and
- controls that do not disappear before a keyboard or assistive-technology user can reach them.
Caption placement should account for names, labels, demonstrations, and other important content near the bottom of the frame. Text should have sufficient contrast against changing backgrounds and should not become unreadable when the video is viewed on a small screen.
HTML caption tracks
HTML provides the <track> element for timed text associated with <video> and <audio>. A basic caption-track implementation may resemble the following:
<video controls>
<source src="/media/filter-replacement.mp4" type="video/mp4">
<track
kind="captions"
src="/media/filter-replacement-en.vtt"
srclang="en"
label="English"
default>
</video>
The presence of a <track> element does not establish that the captions are accurate or that the player exposes them accessibly. Test the rendered player, caption display, keyboard controls, and relevant assistive-technology behavior in the environments the audience is likely to use.
A practical accessible-media production workflow
- Plan before recording.
Identify information that might otherwise be communicated only through gestures, diagrams, or on-screen text. - Retain source materials.
Keep scripts, speaker names, terminology lists, presentation slides, and editable project files. - Work from the final edit.
Caption timing and transcript sequence should correspond to the version that will be published. - Generate or author the captions.
Automatic generation can be a starting point, but the output requires review. - Inventory visual-only information.
Record actions, labels, expressions, charts, and instructions that the soundtrack does not communicate. - Create the appropriate visual alternative.
Depending on the content and requirements, this may be audio description, extended audio description, or a descriptive transcript. - Publish alternatives beside the media.
Use clear labels such as “Read the transcript” or “Play the audio-described version.” - Test the complete experience.
Review the media with keyboard navigation, text resizing, small viewports, and relevant assistive technology. - Update every alternative when the media changes.
New edits can make caption timing, descriptions, transcript text, and links inaccurate.
Review quality, not merely file presence
A media-accessibility review should evaluate whether the alternative preserves meaning accurately and remains usable in context.
Caption review
- Does the text accurately represent the speech?
- Are names, technical terms, numbers, and abbreviations correct?
- Are speaker changes understandable?
- Are meaningful sounds included?
- Does timing follow the corresponding audio?
- Are line breaks and reading speed reasonable?
- Do captions avoid covering essential visual information?
Transcript review
- Does the transcript identify speakers clearly?
- Does it include meaningful sound information?
- If descriptive access is intended, does it include the necessary visual information?
- Can headings, lists, quotations, and other structures be represented with semantic HTML?
- Can readers find the transcript without searching through unrelated downloads?
Audio-description review
- Does the description communicate important visual-only information?
- Does it avoid repeating information already available in the soundtrack?
- Is the wording objective enough to preserve the viewer’s ability to interpret the scene?
- Is the narration timed without obscuring important dialogue or sound?
- Can users locate and activate the described version?
When possible, include people who use captions, transcripts, or audio description in evaluation. Technical checks can identify missing tracks and broken controls, but lived use often reveals timing, wording, and navigation problems that automated review cannot determine.
WCAG and time-based media
The Web Content Accessibility Guidelines 2.2 address prerecorded and live time-based media through several success criteria. At Levels A and AA, these include requirements concerning alternatives for prerecorded audio-only and video-only content, captions for prerecorded and live media, and audio description for prerecorded synchronized media.
WCAG also contains additional Level AAA criteria concerning sign-language interpretation, extended audio description, media alternatives, and live audio-only content. The exact requirement depends on the media, whether it is live or prerecorded, the intended conformance level, and whether a stated exception applies.
The W3C Web Accessibility Initiative provides broader guidance in Making Audio and Video Media Accessible. Publishers working under a law, procurement policy, contract, or organizational standard should verify the requirements that apply to their specific setting. Meeting a technical conformance target also does not remove the need to evaluate whether people can understand and operate the finished media.
Static images within related content require their own text-alternative decisions. See Accessible Images and Alternative Text for guidance on images that do not unfold through time.
Common media-accessibility mistakes
- Treating an automatically generated transcript as finished because words appear on screen.
- Correcting spelling while leaving poor synchronization and unidentified speakers.
- Omitting alarms, music, laughter, or other sounds that change meaning.
- Calling a transcript a complete substitute for synchronized captions in every context.
- Assuming captions communicate visual-only instructions.
- Using translated subtitles that omit meaningful non-speech audio as the only caption track.
- Describing every visual detail rather than identifying what is necessary for understanding.
- Burying a transcript behind an unlabeled icon or ambiguous download link.
- Providing accessible media files inside a player that cannot be operated by keyboard.
- Publishing a revised video with stale captions, descriptions, or transcript timestamps.
- Archiving unreviewed live captions without correcting permanent-recording errors.
Preserve the meaning carried through sound, image, and time
Accessible media is not created by attaching one generic text file to every recording. Captions preserve synchronized audio information. Transcripts provide a readable representation. Descriptive transcripts include essential visual content. Audio description communicates visual meaning during playback.
The appropriate combination becomes clearer when publishers inventory the media itself: what is spoken, what is heard, what is shown, and what someone must understand or do. From there, accessibility becomes a practice of preserving meaning across different modes of access.
art by mary hall https://fine-digital-art.com/