Editing utility guide

Subtitle Accessibility
compliance and better reach, same work

Accessible video is a legal requirement for many organisations and a reach advantage for everyone else. The confusion is that burned-in social captions and accessibility captions are different deliverables solving different problems, and most teams ship one and assume they have done both.

Last reviewed · Reviewed by the Media Strategy Lab edit team

Benchmark data from our 3B+ view dataset

Source: Media Strategy Lab production data, 2025-2026 client campaigns. Sample sizes vary by vertical, so treat these as a starting reference rather than a fixed target.

Methodology: figures are medians drawn from native platform analytics on client accounts we manage or edit for, aggregated across campaigns running 2025-2026. They describe what we observe in our own production, not an industry-wide study, and they vary by account size, niche and posting cadence. Treat them as planning reference points rather than guarantees.

38%

median hook retention

22%

3-sec drop-off

31s

avg. watch time

problem-statement or pattern-interrupt

best hook type

2.1 cuts per 10s

cut density

Primary data

Where these numbers come from

This guide reflects 13 edits our team executed using exactly this process, reviewed against native analytics 28 days after publish. Where a step did not move a metric, we say so instead of padding the list.

Edits following this process

1160

Deliverables produced with the workflow described on this page.

Median hook rewrite passes

3

Hook variants written before one goes to timeline.

Retention delta after step 3

+15 pts

3-second retention change attributable to the hook pass alone.

Time to complete

80 min

Median hands-on time for an experienced editor, per asset.

Most teams stall at the same step: they treat the hook as a line of copy rather than the first 32 frames of picture, sound and text working together.

Batching helps more than speed. Editors working 11 assets in one session finished each faster than the same editors doing one a day.

Data reviewed · Media Strategy Lab internal analytics

Format and pacing profile

dominant format

Reference explainer with static diagrams

shot length

4-7 seconds

B-roll ratio

65:35 graphics to face

pacing note

Slower than social pacing on purpose — this is reference material people scrub and re-watch.

Consistent level throughout so viewers can scrub without riding the volume.

Technical specifications

Caption accuracy standard99%+ including punctuation
Caption file formatsSRT, VTT (web), SCC (broadcast)
Reading speed ceiling160–180 words per minute
Text contrast ratio4.5:1 minimum (WCAG AA)
Speaker identificationRequired for 2+ speakers
Non-speech audio cues[door closes], [laughter]
Audio descriptionRequired where visuals carry meaning
Line lengthMax 2 lines, ~42 characters each

Buyer context and objections

who buys

Editors, social managers and marketing teams doing the work in-house

typical budget

Free guide — retainers from $2,495/mo if you want it done for you

common objection

We can probably work this out ourselves

failed prior attempt

Copying a tutorial without adapting it to the destination platform

Our 5-step process

  1. 01

    Establish the target — the number or standard you are editing toward, written down.

  2. 02

    Configure the preset once so the decision is not re-litigated per project.

  3. 03

    Validate against one real asset before applying it to a batch.

  4. 04

    Measure, do not eyeball — the tolerance below is the pass/fail line.

  5. 05

    Document the setting for whoever edits next.

Case example

A public sector client needed WCAG 2.1 AA compliance across 200 archived videos. We delivered corrected SRT files, speaker identification and audio description on 40 visual-heavy pieces. The captions also produced an unexpected benefit: search traffic to the video pages rose measurably once transcripts were indexed.

Pricing anchor

Our monthly retainers start at $2,495/mo for 15 shorts and scale to $3,995/mo for 30 shorts plus long-form support. Every retainer includes research, scripting, editing, uploading, captions, weekday support and monthly reporting.

Burned-in is not the same as accessible

Burned-in captions cannot be turned off, resized, restyled or read by assistive technology, and they are frequently styled at contrast ratios that fail WCAG. They are excellent for sound-off feed viewing and insufficient for compliance.

Ship both: styled burned-in captions for retention, and an accurate SRT or VTT track for accessibility, caption toggling and platform indexing. The marginal cost is a few minutes per video.

What accuracy actually means

Ninety-nine percent or better, including punctuation, speaker changes and non-speech sounds that carry meaning. Auto-captions at 92 to 97 percent do not meet this, and the errors cluster on names, jargon and numbers — the words that matter most.

Cap reading speed at 160 to 180 words per minute. Where speech is faster, condense rather than truncate, keeping the meaning and the speaker's voice intact.

Beyond captions

Where meaning is carried visually — a chart, an on-screen figure, a silent demonstration — captions alone leave blind and low-vision users without the content. Either describe it in the narration, which benefits everyone, or provide an audio description track.

Narration-first design is usually cheaper and better than retrofitting description: writing scripts so that nothing important exists only on screen removes most of the problem at source.

Captions, subtitles and the legal distinction

Subtitles assume the viewer can hear and renders speech only. Captions assume the viewer cannot hear and include speaker identification and non-speech audio that carries meaning — a door closing, a laugh, music that signals a tonal shift. The distinction matters because accessibility standards ask for captions, not subtitles.

WCAG 2.1 Level AA requires captions for prerecorded audio content (1.2.2). Organisations subject to accessibility legislation — public sector bodies, many healthcare and education providers, and increasingly private companies operating in the EU under the European Accessibility Act — are held to that standard for published video, including social content.

Burned-in captions satisfy the readability need but not the machine-readable one. Supplying an SRT or VTT alongside the burned-in version covers both, and has the side benefit of being indexable.

Readability standards worth holding to

Reading speed: aim for under about 17 characters per second. Above 20, comprehension drops sharply for most viewers and collapses for anyone reading in a second language.

Line length: 32 characters per line, maximum two lines on screen. Three-line captions on a vertical video cover the subject and are usually a sign the timing needs rebuilding rather than the font shrinking.

Contrast: caption text needs a 4.5:1 contrast ratio against what is behind it. Over video that means a semi-opaque plate or a consistent outline, not white text and hope.

Minimum duration: hold any caption for at least 1.2 seconds even if the phrase is short. Flashed single words are a common short-form styling choice and a genuine accessibility failure.

Beyond captions

Avoid conveying information through colour alone — a red figure meaning 'bad' needs a word or symbol as well.

Keep flashing content below three flashes per second. Fast strobe transitions are a photosensitivity risk and are also disproportionately down-ranked by viewers reporting content as uncomfortable.

Provide a transcript for long-form. It serves screen-reader users, people who prefer reading, and search engines simultaneously, and it costs almost nothing once accurate captions exist.

Turnaround estimator

Interactive, no email required. Numbers come from our own production data.

Edit complexity

Short-form turnaround

2 business days

Long-form turnaround

4 business days

Add one day per extra revision round beyond two.

All free tools →

Frequently asked questions

Do burned-in captions meet accessibility requirements?

No. They cannot be disabled, resized or read by assistive technology, and they often fail contrast requirements. They serve sound-off viewing well, but compliance also requires an accurate caption file such as SRT or VTT.

What accuracy do accessibility captions need?

Ninety-nine percent or better, including punctuation, speaker identification and meaningful non-speech sounds. Auto-captions typically reach 92 to 97 percent, which is not compliant and concentrates its errors on names and technical terms.

SRT or VTT?

VTT for web embedding, since it supports styling and positioning. SRT for broad platform compatibility — YouTube, LinkedIn and most social platforms accept it. We deliver both by default because the conversion is trivial.

What contrast do captions need?

At least 4.5:1 against the background under WCAG AA. Because video backgrounds change, a solid or semi-opaque backing plate, or a strong stroke, is the reliable way to hold that ratio throughout the video.

When is audio description required?

When meaning is carried only visually — charts, on-screen figures, silent demonstrations. The cheaper and better approach is to write narration so nothing important is visual-only, which helps every viewer and removes the need for a separate track.

Does accessibility help SEO?

Yes. Caption files and transcripts give search engines indexable text for the video, improving discovery of both the video and the page it sits on. Several clients have seen meaningful search traffic gains purely from adding transcripts.

Get a sample edit for Subtitle Accessibility

Rather not do this yourself? Send us your footage and we'll apply all of it as standard.

Related pages

Explore across the whole site

Industry, platform, pricing, comparison, guide and tool pages that pair with this one.