Captioning / Transcripts
Captions are synchronised text that displays spoken dialogue, sound effects, and other audio information during video playback. Transcripts are full-text records of audio or video content. Both are WCAG Level A requirements for pre-recorded multimedia. Captions serve people who are deaf or hard of hearing, non-native speakers, and anyone watching without sound.
§ 1 Definition
Captions and transcripts are text alternatives for audio content in multimedia presentations. Captions are synchronised with the video timeline and display spoken words, speaker identification, and important non-speech audio (such as [door creaks] or [phone rings]). They can be open (always visible) or closed (user-toggleable via the CC button). WCAG 2.1 requires captions for all pre-recorded synchronised media at Level A (1.2.2 Captions) and live audio at Level AA (1.2.4 Captions Live). Transcripts are a full text record of the audio content plus, for video, descriptions of the visual content. Transcripts are required at Level A (1.2.3 Audio Description or Media Alternative) for pre-recorded video. Automated captioning tools (YouTube, Zoom, Otter.ai) provide a starting point but are not accurate enough for WCAG compliance. Every automated caption set must be reviewed and corrected manually for accuracy, timing, and speaker identification.
§ 2 Captions vs subtitles: what is the difference
The terms are often used interchangeably, but they serve different purposes. Subtitles assume the viewer can hear the audio but do not speak the language, so they translate dialogue only. Captions (specifically closed captions) are designed for viewers who cannot hear the audio, so they include dialogue, sound effects, music cues, and speaker identification, all synchronised to the video timeline. Captions should use the original language and should not be translated. For multilingual audiences, provide both captions and translated subtitles as separate tracks.
§ 3 WCAG requirements and timing
For pre-recorded video with audio: captions must be provided (1.2.2, Level A). For live video (webinars, live streams): captions must be provided (1.2.4, Level AA). For pre-recorded audio-only content with no video (podcasts): a transcript is the primary accessibility method. The transcript must include all spoken content and identify speakers. For pre-recorded video: in addition to captions, an audio description or a full text transcript describing the visual content is required (1.2.3, Level A). Audio description describes what is happening visually (actions, expressions, scene changes) during natural pauses in the dialogue. Full transcripts for video combine all dialogue, sound effects, and visual descriptions into one document.
§ 4 Common questions
- Q. Are auto-captions (YouTube, Zoom) good enough for WCAG?
- A. No. Automated captions typically achieve 80-90 percent accuracy, which is not enough for WCAG compliance. Errors are especially common with technical terms, names, accents, and overlapping speech. You must manually review and correct every caption set. Many video platforms allow you to upload corrected caption files (SRT, VTT) after editing.
- Q. Do I need transcripts for every podcast?
- A. Yes, if you host the podcast on your website or make it available to the public. A full transcript of the spoken content serves deaf and hard of hearing audiences, supports search engine indexing, and allows users to scan and find specific topics. Providing a transcript is also a WCAG Level A requirement for audio-only content.
- Captions are synchronised text for video, required at WCAG Level A (pre-recorded) and AA (live).
- Transcripts are full-text records of audio content, required for pre-recorded audio and video.
- Auto-captions are a starting point, not a compliance solution. Review and correct every caption file.
- Captions include sound effects and speaker IDs, not just dialogue.
- Transcripts improve SEO and content discoverability in addition to accessibility.
Atomic Glue can help you implement accessible video players with proper caption and transcript support. We test caption timing, AC3 compliance, and transcript navigability as part of our Web Development services. Get in touch to discuss multimedia accessibility.
Get in touchCaptions are synchronised text that displays spoken dialogue, sound effects, and other audio information during video playback. Transcripts are full-text records of audio or video content. Both are WCAG Level A requirements for pre-recorded multimedia. Captions serve people who are deaf or hard of hearing, non-native speakers, and anyone watching without sound.
Category: Accessibility (also: General, Content Strategy)
Author: Atomic Glue Team
## Definition
Captions and transcripts are text alternatives for audio content in multimedia presentations. Captions are synchronised with the video timeline and display spoken words, speaker identification, and important non-speech audio (such as [door creaks] or [phone rings]). They can be open (always visible) or closed (user-toggleable via the CC button). WCAG 2.1 requires captions for all pre-recorded synchronised media at Level A (1.2.2 Captions) and live audio at Level AA (1.2.4 Captions Live). Transcripts are a full text record of the audio content plus, for video, descriptions of the visual content. Transcripts are required at Level A (1.2.3 Audio Description or Media Alternative) for pre-recorded video. Automated captioning tools (YouTube, Zoom, Otter.ai) provide a starting point but are not accurate enough for WCAG compliance. Every automated caption set must be reviewed and corrected manually for accuracy, timing, and speaker identification.
## Captions vs subtitles: what is the difference
The terms are often used interchangeably, but they serve different purposes. Subtitles assume the viewer can hear the audio but do not speak the language, so they translate dialogue only. Captions (specifically closed captions) are designed for viewers who cannot hear the audio, so they include dialogue, sound effects, music cues, and speaker identification, all synchronised to the video timeline. Captions should use the original language and should not be translated. For multilingual audiences, provide both captions and translated subtitles as separate tracks.
## WCAG requirements and timing
For pre-recorded video with audio: captions must be provided (1.2.2, Level A). For live video (webinars, live streams): captions must be provided (1.2.4, Level AA). For pre-recorded audio-only content with no video (podcasts): a transcript is the primary accessibility method. The transcript must include all spoken content and identify speakers. For pre-recorded video: in addition to captions, an audio description or a full text transcript describing the visual content is required (1.2.3, Level A). Audio description describes what is happening visually (actions, expressions, scene changes) during natural pauses in the dialogue. Full transcripts for video combine all dialogue, sound effects, and visual descriptions into one document.
## Common questions
Q: Are auto-captions (YouTube, Zoom) good enough for WCAG?
A: No. Automated captions typically achieve 80-90 percent accuracy, which is not enough for WCAG compliance. Errors are especially common with technical terms, names, accents, and overlapping speech. You must manually review and correct every caption set. Many video platforms allow you to upload corrected caption files (SRT, VTT) after editing.
Q: Do I need transcripts for every podcast?
A: Yes, if you host the podcast on your website or make it available to the public. A full transcript of the spoken content serves deaf and hard of hearing audiences, supports search engine indexing, and allows users to scan and find specific topics. Providing a transcript is also a WCAG Level A requirement for audio-only content.
## Key takeaways
- Captions are synchronised text for video, required at WCAG Level A (pre-recorded) and AA (live).
- Transcripts are full-text records of audio content, required for pre-recorded audio and video.
- Auto-captions are a starting point, not a compliance solution. Review and correct every caption file.
- Captions include sound effects and speaker IDs, not just dialogue.
- Transcripts improve SEO and content discoverability in addition to accessibility.
## Related entries
- [Accessibility (a11y)](atomicglue.co/glossary/accessibility)
- [WCAG (2.1, 2.2)](atomicglue.co/glossary/wcag)
Last updated July 2026. Permalink: atomicglue.co/glossary/captioning-transcripts