Captions, audio description and CART are separate purchases
The line item says “captioning services.” One vendor, one purchase order, one renewal date. Underneath it sit four obligations from different rules, attached to different units of delivery, failing separately: captions on prerecorded media, captions on live media, audio description on prerecorded video, and real-time human transcription arranged for a person who asked for it. Buy them as one thing and you will discharge one of them.
The clock is the ADA Title II web and mobile app rule. 28 CFR 35.200(b)(1) provides that “Beginning April 26, 2027,” a public entity with a population of 50,000 or more “shall ensure that the web content and mobile apps that the public entity provides or makes available, directly or through contractual, licensing, or other arrangements, comply with Level A and Level AA success criteria and conformance requirements specified in WCAG 2.1.” Paragraph (b)(2) sets April 26, 2028 for smaller entities and special district governments. Those are not the dates the 2024 final rule published. They come from an interim final rule at 91 FR 20902, published and effective April 20, 2026, whose comment period closed on June 22, 2026, so a final rule could still move them. It is worth reading before you re-baseline anything.
Note the phrase in the middle of it: “directly or through contractual, licensing, or other arrangements.” A vendor’s captioning failure is the public entity’s failure, which makes this a drafting problem, not a vendor selection problem.
Four obligations, four places in the rulebook
WCAG 2.1 splits time-based media across nine success criteria. Five sit inside the Level A and AA set Title II incorporates, and four of those five are purchases. 1.2.2 Captions (Prerecorded), Level A: “Captions are provided for all prerecorded audio content in synchronized media, except when the media is a media alternative for text and is clearly labeled as such.” 1.2.4 Captions (Live), Level AA: “Captions are provided for all live audio content in synchronized media.” 1.2.3, Level A, lets an author choose audio description or a text alternative; 1.2.5 Audio Description (Prerecorded), Level AA, closes that choice: “Audio description is provided for all prerecorded video content in synchronized media.” The fifth, 1.2.1, is where audio-only and video-only content lands. Cite the dated text, since 28 CFR 35.200(b)(3) incorporates the WCAG 2.1 Recommendation of 05 June 2018, not the undated address, which now serves a 2025 revision.
Federal agencies arrive by another route: the Revised 508 Standards at 36 CFR part 1194, Appendix A, E205.4, require electronic content to conform to Level A and AA “in WCAG 2.0.” For media that difference does not bite, because criteria 1.2.1 through 1.2.9 came into WCAG 2.1 from WCAG 2.0 unchanged. Recipients of HHS financial assistance sit on a third clock, 45 CFR 84.84(b): May 11, 2027 with fifteen or more employees, May 10, 2028 with fewer.

View the data as a table
| Prerecorded audio, synced media | Live audio, synced media | Prerecorded video, synced media | Live audio-only stream | |
|---|---|---|---|---|
| WCAG criterion | 1.2.2 Captions (Prerecorded) | 1.2.4 Captions (Live) | 1.2.5 Audio Description (Prerecorded) | 1.2.9 |
| Conformance level | Level A | Level AA | Level AA | Level AAA |
| In the incorporated set | Yes | Yes | Yes | No, it sits above it |
| What you buy | Captions on the recording | Live captions | Audio description | Nothing at Level A or AA |
Three cases save money. Prerecorded audio-only content is not a captions purchase: 1.2.1 asks for “an alternative for time-based media,” which for a podcast is a transcript. A live stream with no video is not synchronized media, so 1.2.4 does not reach it, and audio-only live sits at 1.2.9, Level AAA, above the incorporated set. And audio description is not needed where the audio already carries the picture: “Where all of the video information is already provided in existing audio, no additional audio description is necessary.”
Two exceptions decide how much of the library you are pricing
Before pricing a single item, run the library against 28 CFR 35.201. Paragraph (a) takes archived web content out of 35.200 altogether, and 28 CFR 35.104 defines it in four parts, all of which have to hold: created before the entity’s compliance date, “retained exclusively for reference, research, or recordkeeping,” “not altered or updated after the date of archiving,” and “organized and stored in a dedicated area or areas clearly identified as being archived.” A three-year-old meeting recording sitting in the same folder as this month’s is not archived within that definition. Paragraph (c) excepts “content posted by a third party, unless the third party is posting due to contractual, licensing, or other arrangements with the public entity,” so the terms on which a platform hosts your stream decide whether it counts. Segregating an archive is a procurement decision before it is a captioning decision, and it sets the number of items you pay to caption.
What counts as live decides which contract an item falls under
WCAG defines “live” as “information captured from a real-world event and transmitted to the receiver with no more than a broadcast delay,” and “prerecorded” as “information that is not live.” Live audio inside synchronized media falls under 1.2.4 at Level AA; a recording of the same event falls under 1.2.2 at Level A, with a different vendor process behind it.
The W3C’s captions resource names the seam: “If you have live captions and you post a recording, you will probably need to do minor editing for accuracy.” Two deliverables. Buy the first, publish the second, and the file on your website is an unedited real-time transcript: what satisfied 1.2.4 has not satisfied 1.2.2. The FCC agrees, in voluntary best practices at 47 CFR 79.1(k)(2)(xvi): real-time captioning is “for live and near-live programming, and not for prerecorded programming.”
No rule sets a number for caption quality
WCAG 1.2.2 and 1.2.4 say captions are provided. No accuracy figure, no latency figure, no completeness test. The Department of Justice was asked to fix that and declined, at 89 FR 31320, 31359: the Department “does not believe it is prudent to prescribe captioning requirements beyond the WCAG 2.1 Level AA requirements, whether by specifying a numerical accuracy standard, a method of captioning that public entities must use to satisfy this success criterion, or other measures.”
The closest DOJ comes to a quality statement is on its web guidance page, which predates the rule: videos “can be made accessible by including synchronized captions that are accurate and identify any speakers in the video.” Three attributes, no threshold on any of them, and no test.
On automatic captions the preamble goes this far and no further: the Department “recognizes commenters’ concerns that automatic captions are currently not sufficiently accurate in many contexts,” and notes “that informal guidance from W3C provides that automatic captions are not sufficient on their own unless they are confirmed to be fully accurate.” The W3C material is blunter: “missing just one word such as ‘not’ can make the captions contradict the actual audio content.” Who confirms an automatic file is accurate, against what, and how that is evidenced, is a role your contract has to invent.
The only caption specification in federal law binds someone else
One federal rule defines accurate, synchronous, complete and appropriately placed captions in regulatory text: 47 CFR 79.1, “Closed captioning of televised video programming,” reaching programming “distributed and exhibited for residential use.” It binds video programming distributors and programmers, not a city buying captions for a council livestream, and DOJ declined to impose anything like it on public entities. Borrowed by analogy as drafting language, it gives a buyer four defined terms that WCAG does not.
Paragraph (j)(2) sets the frame: “Captioning shall be accurate, synchronous, complete, and appropriately placed as those terms are defined herein.” Accuracy, at (j)(2)(i), is verbatim with two limits written into it. Captioning “shall match the spoken words (or song lyrics when provided on the audio track) in their original language (English or Spanish), in the order spoken, without substituting words for proper names and places, and without paraphrasing, except to the extent that paraphrasing is necessary to resolve any time constraints.” It must also supply “nonverbal information that is not observable, such as the identity of speakers, the existence of music (whether or not there are also lyrics to be captioned), sound effects, and audience reaction, to the greatest extent possible, given the nature of the program.” Carry the qualifiers across with the duty. A clause that bans paraphrase outright, or demands every nonverbal cue, is stricter than the rule it came from.
Synchronicity, at (j)(2)(ii), is that captions “begin to appear at the time that the corresponding speech or sounds begin and end approximately when the speech or sounds end,” at “a speed that permits them to be read by viewers.” Completeness, at (j)(2)(iii): “Captioning shall run from the beginning to the end of the program, to the fullest extent possible.” Placement, at (j)(2)(iv), is that captioning “shall not block other important visual content on the screen, including, but not limited to, character faces, featured text (e.g., weather or other news updates, graphics and credits), and other information that is essential to understanding a program’s content.” The rule carries on past that point, into font size, line overlap and captions running off the edge of the screen.

View the data as a list
47 CFR 79.1(j)(2): Binds TV distributors and programmers, not a city buying captions
- Accurate: Verbatim, proper names kept, paraphrase only where timing forces it
- Synchronous: Begins and ends with the speech, at a speed that permits reading
- Complete: Beginning to the end of the program, to the fullest extent possible
- Appropriately placed: No blocking of faces, featured text, graphics or credits
Now the number, and the label it needs. Paragraph (k)(2)(iv) tells real-time captioning vendors to treat accuracy as “the percentage of correct words out of total words in the program”: “For example, 7,000 total words in the program minus 70 errors equals 6,930 correct words captioned, divided by 7,000 total words in the program equals 0.99 or 99% accuracy.” Paragraph (k)(2)(v) counts as errors “mistranslated words, incorrect words, misspelled words, missing words, and incorrect punctuation that impedes comprehension and misinformation.”
Nobody is required by law to hit 99 percent. The Commission said in the adopting release at 79 FR 17911, 17925: “First, the Best Practices are voluntary,” and the standards it did adopt are “qualitative rather than quantitative.” One route pulls the number closer: under (j)(1)(i)(B) a video programmer may certify that it “has adopted and follows the Best Practices set forth in paragraph (k)(1),” which at (k)(1)(i)(A) calls for performance requirements “comparable to those described in paragraphs (k)(2), (k)(3) and (k)(4).” The only percentage attached to caption quality in federal law sits in voluntary guidance, addressed to another industry, for live work alone. The mandatory percentages in 79.1, “100% of new, nonexempt” and “75% of pre-rule, nonexempt” programming, measure how much gets captioned, not how good it is.
Live work alone matters. The offline captioning best practices at (k)(4) set a stricter bar with no percentage in it: “Ensure offline captions are verbatim,” “Ensure offline captions are error-free,” and a “complete textual representation of the audio, including speaker identification and non-speech information.” The FCC treats live and prerecorded captioning as two specifications, which is this article’s thesis written into a federal rule by an agency that had to price the difference.
The clauses that text hands you, and two duties it hands back
None of this binds a public entity. It is the only federal text here that reads like a captioning contract, which is what makes it worth borrowing. Paragraph (k)(1), the video programmer best practices, asks buyers to put into vendor agreements “performance requirements designed to promote the creation of high quality closed captions,” “a means of verifying compliance with such performance requirements, such as through periodic spot checks,” and training provisions. Two clauses on the vendor’s side of the same rule, at (k)(2)(x) and (k)(2)(xi), are worth mirroring into a live spec, because a per-item purchase order cannot express staffing: captioners “qualified for the type and difficulty level of the programs to which they are assigned,” and “a system that verifies captioners are prepared and in position prior to a scheduled assignment.”
Two clauses point back at the buyer, and that is the part left out. At (k)(1)(ii)(A) the programmer provides preparation materials “to the extent available”: “show scripts, lists of proper names (people and places), and song lyrics.” At (B) it makes “commercially reasonable efforts” to supply “a high quality program audio signal to promote accurate transcription and minimize latency.” An agenda with every council member’s name spelled correctly, sent the day before, and a clean audio feed are two cheap quality interventions, and both sit on the buyer’s side.
Audio description is required at Level AA and absent from the guidance
Here is a measurement rather than an impression. The Title II final rule preamble runs 77 pages, 89 FR 31320 to 31396. The string “caption” appears 82 times. The string “audio description” appears zero times, and the 21 success criteria the preamble names by number include none of 1.2.3, 1.2.5 or 1.2.7. DOJ’s web guidance page contains no occurrence of the phrase either. Yet 1.2.5 is Level AA, and 28 CFR 35.200(b) incorporates Level AA. Audio description is required by the same sentence that requires captions, and every word of federal guidance about media under Title II is about the other one.
Two distinctions keep the purchase from being doubled or oversold. First, 1.2.3 and 1.2.5 are one purchase, not two. At Level A an author may choose audio description or a full text alternative; at Level AA audio description is, in the words of W3C’s Understanding document for 1.2.5, explanatory material published alongside WCAG 2.2 rather than part of the standard, “a requirement already met if they chose that alternative for 1.2.3, otherwise an additional requirement.”
Second, standard and extended audio description are different products at different levels. WCAG defines audio description as “narration added to the soundtrack to describe important visual details,” added in the standard kind “during existing pauses in dialogue.” Extended audio description “is added to an audiovisual presentation by pausing the video,” and sits at 1.2.7, Level AAA, above the Title II and Section 508 standard. The FCC’s definition at 47 CFR 79.3(a)(3) covers only the standard kind: “The insertion of audio narrated descriptions of a television program’s key visual elements into natural pauses between the program’s dialogue.”
Description comes in three delivery forms, which are three line items: integrated description written into the speakers’ script, an alternative described video, and a “separate file” that “must be supported by the media player.” W3C puts the first at zero, with the scope limiter that belongs in the quote: “For some types of video (such as some training videos), description of the visual information can be seamlessly integrated by the speakers as the video is planned and created, and you don’t need separate description, thus there is no additional cost.” That is a statement about how a video was planned, not a price for description as a product.
Live audio description sits outside WCAG. The W3C’s description resource says of live content: “Description is needed to provide the important visual information to people who are blind. Description is not required to meet WCAG.” Anyone who needs it at a live event needs it through the effective communication rules, which is the next purchase.
CART is the fourth purchase, and WCAG does not scope it
CART is not named in Title II by that name. Across the full text of 28 CFR part 35, “Communication Access Realtime Translation” and “CART” appear zero times. What the regulation names, in the definition of “auxiliary aids and services” at 28 CFR 35.104, is “real-time computer-aided transcription services” and “open and closed captioning, including real-time captioning.” That is the regime CART is bought under, and it is not the web rule. The duty is at 28 CFR 35.160: communications must be “as effective as communications with others,” and in choosing the aid “a public entity shall give primary consideration to the requests of individuals with disabilities.”
That is why a CART contract cannot be scoped from a menu the entity picks alone. The aid “will vary in accordance with the method of communication used by the individual; the nature, length, and complexity of the communication involved; and the context,” and must protect “the privacy and independence of the individual with a disability.” What you are buying, in the W3C’s words, is that “live captions are usually done by professional real-time captioners or Communication Access Realtime Translation (CART) providers”: a scheduled, qualified human, booked in advance. No rule cited here sets a rate, and this article does not state one.
Two boundaries. WCAG 1.2.4 does not push captioning onto a conferencing product. W3C’s Understanding document for that criterion, explanatory rather than normative and cited by DOJ in the preamble at footnote 119, says the criterion “was intended to apply to broadcast of synchronized media and is not intended to require that two-way multimedia calls between two or more individuals through web apps must be captioned,” with responsibility on the callers or the host. And the two routes can diverge for one event. A streamed council meeting is covered by 1.2.4 whether or not anyone asks; the same meeting attended by someone who requests CART is covered by 35.160, where the request gets primary consideration. Two obligations, possibly two vendors, and nothing in the rules cited here reconciles them.
DOJ heard the lead-time argument and answered it without building a runway. Commenters noted “the lead time necessary to reserve those services” (89 FR 31358). The Department replied that the compliance dates “will give public entities sufficient time to locate captioning resources and implement or enhance processes to ensure they can get captioning services when needed,” then set “a uniform compliance date for all success criteria in subpart H” (89 FR 31359).

View the data as a table
| Live captions | CART | |
|---|---|---|
| Rule it sits under | WCAG 1.2.4, inside the web rule | 28 CFR 35.160, not the web rule |
| What triggers it | Streaming the meeting, whether or not anyone asks | An individual requests it |
| Who scopes it | The criterion, not any request | The request, given primary consideration |
| What you buy | Captions on the broadcast stream | A scheduled, qualified human, booked in advance |
The player is a product requirement, not a content one
A caption file no control can switch on is not accessible media, and here the Revised 508 Standards do something WCAG does not: they bind the hardware and the software. Four provisions, published on the Access Board’s ICT page and codified in 36 CFR part 1194, Appendix C. Provision 413.1 requires ICT that “displays or processes video with synchronized audio” to provide “closed caption processing technology” that decodes and displays captions, or passes caption data through; 414.1 says the same for “audio description processing technology.” Provision 415.1 covers hardware controls: operable parts for caption selection wherever there are operable parts for volume, and for audio description wherever there are operable parts for program selection, excepting personal devices where “captions and audio descriptions can be enabled through system-wide platform settings.” Provision 503.4 is the software analogue, and it holds the one test a buyer can run inside a demo: caption and description controls must sit “at the same menu level as the user controls for volume or program selection.”
Section 508 reaches vendors through the acquisition regulation. FAR 39.203(a) requires that “acquisitions for ICT supplies and services shall meet the applicable ICT accessibility standards at 36 CFR 1194.1” unless an exception applies, and 39.203(c) makes each task or delivery order its own compliance event. If nothing on the market conforms, 39.205(a)(3) lowers the requirement to a floor rather than removing it: procure what “best meets the ICT accessibility standards consistent with the agency’s needs,” then provide access by an alternative means. Our Section 508 compliance work starts there, because “no conforming product exists” is a determination with required contents, not a conclusion.

View the data as a list
Player conformance: A caption file no control can switch on is not accessible media
- 413.1 Caption processing: Decodes and displays captions, or passes caption data through
- 414.1 Description processing: The same duty for audio description processing technology
- 415.1 Hardware controls: Caption selection wherever there are operable parts for volume
- 503.4 Software controls: Caption and description controls at the volume menu level
What is not settled
No published decision applying 28 CFR 35.200 to captions turned up in the research for this article. The compliance dates are still ahead, so no public entity is yet measured against the rule.
There is no de minimis provision for web content matching the FCC’s. The FCC carve-out at 79.1(j)(3) weighs “the type of failure, the reason for the failure, whether the failure was one-time or continuing, the degree to which the program was understandable despite the errors.” Title II instead has 28 CFR 35.205, excusing noncompliance with “such a minimal impact on access” that participation keeps “substantially equivalent timeliness, privacy, independence, and ease of use.” How that applies to captions has not been decided.
Whether an automatic caption file can satisfy 1.2.4 is unresolved by regulation. DOJ declined to say and pointed at W3C material requiring confirmation of full accuracy, without defining who confirms, against what sample, or how it is recorded. The preamble cites that page as of July 14, 2022, and the page’s own source file carries the same date, so there is no gap between what DOJ read and what is published today.
What to put in the contract
- Split the line item into four, after you scope the library. Items excepted by 35.201 are not priced at all. Then: prerecorded captions per item of synchronized media, with a transcript rather than captions for anything audio-only. Audio description per item of prerecorded video whose visuals the audio does not carry. Live captions per scheduled event, staffed. Post-event caption files as a deliverable separate from the live session.
- Write the quality specification, since no rule will write it for you. Borrow 79.1(j)(2)(i) through (iv) as drafting language with the label attached: an FCC rule for television programmers, not law that binds you. Verbatim except where timing forces a paraphrase, proper nouns correct, speaker identification present, non-speech audio notated, a stated maximum lag, whole asset start to finish, no occlusion of faces or on-screen text. If you want a percentage, say where it came from and that you elected it.
- Buy the live service as a service. Qualified captioner for the subject matter, in position before start, plus your own duties: preparation materials in advance, a clean audio feed, and an issue log recording “date, time of day, program title, and description of the issue.”
- Test the player before the content. Caption control at the same menu level as volume. Audio description selectable where program selection exists. If you buy description as a separate timed file, confirm the player supports it before commissioning any.
For a government entity or a college or school system with a media library and a 2027 date, live captioning has the longest lead time and the least flexible unit, and audio description is the obligation no federal guidance has told anyone to look at, which is why it is the one least likely to be in your current scope.
One boundary. This is procurement and testing guidance. Whether a specific program, service or activity carries a given obligation, how the undue burden defense at 28 CFR 35.204 applies to your budget, and what to concede in a complaint, are questions for your counsel.