Accessible videos allow
all students and staff to fully engage with course content and campus
resources, including people who are deaf or hard of hearing, blind or
low-vision, or who use text-based formats to support learning. Accessible design also benefits everyone: captions help in noisy settings, transcripts support skimming and study, and clear controls make playback easier on all devices.
To meet accessibility standards, every video must provide:
Captions: accurate text of spoken words and meaningful sounds
Transcripts: a complete text version of the content, useful for offline reading, search, and braille access
Audio Description: narration of essential visual details that are not spoken aloud
Accessible Player Controls: keyboard controllable actions for play, pause, volume, and track selection. Videos should also not use auto-play when a page loads, users should choose to select play.
Start with a script or storyboard that includes what will be said and what will be shown on-screen—graphics, charts, screen text, etc. That helps you identify what needs description, what visuals are essential, what might confuse a user relying on audio description. (W3C “Planning Audio & Video Media”)
Keep segments short and focused
Choose MP4 output. Record at a resolution that fits your platform
Leave empty space in the layout if you will include an interpreter window
Use clear audio: good mic, quiet environment, speak slowly and distinctly. Avoid overlapping speech (e.g., avoid speaking while animations or transitions are happening).
Make sure visuals are accessible:
High contrast between text/graphics and background.
Large, legible fonts.
Avoid cluttered slides. If you must present a lot, break it into multiple visuals or slides.
Leave natural pauses after complex information for processing, and to give room for later description or narration of visuals.
If
you use slides or screen capture, ensure any text on screen is also
read aloud by the speaker. Describe charts, graphs, or complex visuals
so users relying on audio descriptions/described video still understand
what's going on.
Always provide captions. Auto-captions are a starting point; review and correct them.
Ensure captions are accurate and complete.
Identify speakers and include meaningful non-speech sounds (e.g., [music], [laughter], [applause]).
Use a sans-serif, readable font; large enough size; and consistent placement (bottom center, unless blocking visuals).
Provide transcripts in an accessible format (HTML page, or Word file).
Basic transcripts (spoken words only) are acceptable for audio-only content.
Descriptive transcripts (words + key visual details) are needed when there is no audio description track, and descriptions of visual content are missing within the video's narration.
Place transcripts near the video or provide a clear, accessible link.
Required if essential visual information is not already described in the narration.
Can be built into the main narration or provided as a separate audio description track.
Keep descriptions concise, objective, and inserted into natural pauses so they don’t overlap with speech.
If a separate audio track isn’t feasible (Canvas does not support this), provide a descriptive transcript as an alternative.
Video without audio: provide an audio description track and a descriptive transcript
Audio-only content: provide a full transcript. For live audio events, provide live captions when feasible
Live presentations: provide real-time captioning if possible; after the event, supply edited captions and transcripts
Are auto-captions enough? No, auto-captions are a helpful first draft. Accuracy varies by audio quality, accents, jargon, and live streaming. Human review of captions is required.
Burned-in (Open) Captions: Permanently embedded into the video (can’t be turned off).
Not machine-readable → not accessible for screen readers or captioning software.
If using burned-in captions, you would also have to provide a closed caption file to meet accessibility standards.
Closed Captions: Appear as a toggle-able text track (e.g., via a CC button). Typically include:
Spoken dialogue
Speaker identification
Non-speech elements (like [laughter], [door slams], or music cues)
Machine-readable → compatible with screen readers, search, and accessibility tools.
Best for accessibility
Subtitles: Typically
used to translate dialogue into another language. Usually include
dialogue only; non-speech elements (like [sighs] or [music]) are often
excluded. Can be:
Open (burned-in) → permanently on
Closed (toggleable) → like CC, but often without sound cues
Subtitles can be similar to burned in captions visually but are in a language different than the audio track file of the multimedia
Works when video will not play or internet is slow
Many people prefer reading or need print
Supports users who read with refreshable braille
Great for skimming, quoting, and search
Choose one and label it clearly:
A download link to a .docx or .txt file under the video
An accordion that expands below the video, with an ARIA label such as "Click Here for the [Video Name] Transcript"
A separate HTML page with the content as plain text and the page is linked under the video
All spoken words (excluding filler words like "um" and false starts)
Use a readable sans serif font at least 11pt in size
List the speaker's name followed by a colon. (Ex: Jim:)
List the speakers name each time there is a transition of who is speaking
Decide on retaining or removing timestamps based on user needs
Timestamps
allow readers to match text with audio/video. It’s recommended to
include at least limited time stamps at heading levels and/or when
speakers transition
Ex: Heading [2:00]
Ex: Jim [2:00]:
Include on-screen information (required for content not included in audio but best practice for all on-screen content)
On-Screen: [Text content inside brackets]
(Sounds outside of speech, like an alarm sound, should be contained within parentheses)
If using headings in your transcript:
Consider inserting a heading when a new topic or interview question is introduced
Ensure you’re following existing heading accessibility guidance (logical order -- H1, H2, H3, etc.)
One solution is to put the video link in ChatGPT and ask it to provide the text of the transcript only (Note: the free version of ChatGPT can only handle videos less than 30 minutes long).
Zoom provides a VTT file that can be the text of the transcript
Many automated files will be in single line sentences, so it is best practice to edit this format into readable paragraph length text as appropriate.
We recommend using AI to assist and suggest the below as a prompt option:
"Review the attached timestamped transcript and update capitalization, punctuation, remove repeats of names of who is talking, so that the name shows once each time they speak and not again until someone new is talking. Speakers names should be bold. Remove the timestamps for this transcript but do not change any language or phrasing. Describe non-speech sounds that are important for context."
Yes – If you are linking to the video and all users travel to YouTube to view it, they can access the transcript file directly on YouTube.
No
– If you embed the video into Canvas or another platform, the
transcript is not automatically visible to viewers. In that case, you
need to:
Ensure captions are turned on and accurate, and
Provide the transcript separately (either posted under the video or linked as a text file).
Here are your options:
If the video has captions but no transcript:
A transcript can often be pulled from the caption file that’s embedded in the video.
You can often locate and export the caption file by looking at the video’s HTML code. Reach out to our IT Department for help with this.
If the video does not have captions:
The video will not be accessible in its current form. You’ll need to either:
Find an alternative version of the video that already has captions, or
Reach out to the owner of the video and request they provide an accessible version
When is audio description needed?
Add it when visuals carry meaning that the audio does not include. Examples: labeled diagrams, on-screen steps, or animations.
Best practice
Write your script so the speaker describes essential visuals out loud. This often removes the need for a separate description track. If gaps remain, add a description track or provide a descriptive transcript.
How to decide if you need a separate track or transcript
Ask yourself:
Is there on-screen text, a change of speaker, or other essential information not included in the audio?
If yes:
First check if the audio can be updated (narration or script edits) to include the information.
If yes → update the audio. No separate description is needed.
If no → add descriptions into the transcript (preferred), or add in a described video.
If no: no audio description is needed
How to handle audio descriptions in Canvas
For
many course videos, it’s often easiest to provide a descriptive
transcript that combines the spoken words with descriptions of key
visuals. This can be placed directly below the video, inside a
collapsible <details> section, or as a linked file.
In Canvas, videos cannot host multiple audio tracks. If you make a described version, upload it as a second video (one standard, one described).
Tips from W3C Web Accessibility Initiative (WAI)
Essential visuals include on-screen text, speaker names shown visually, diagrams, and scene changes not covered in narration.
Keep descriptions short, objective, and in the present tense.
Ensure your player supports description tracks or timed text, if you use them.
For more information, see W3C's Audio Description Checklist
Use a player with keyboard-accessible controls for play, pause, volume, and track selection
Turn off auto-play on page load
If you link to YouTube, share how to manage Autoplay in YouTube settings