VPAT ACR

What your ACR has to say about the AI panel you shipped

David LoPresti By David LoPresti August 22, 2026

The questionnaire came back with one line added since last year: send the Accessibility Conformance Report, and make sure it covers the assistant. The assistant is the panel your team shipped this year. It streams its answer a few tokens at a time, sometimes redraws a paragraph it has already rendered, and ends with three follow-up chips whose wording nobody on your team wrote.

Then you open the template and there is nothing in it about any of that. ITI publishes four current editions of the VPAT 2.5Rev template, all released in April 2025: 508, EU, WCAG and INT. Extract the plain text of the 508 edition and the International edition and search each, case-insensitive, for “artificial”, “generative”, “chatbot”, “chat bot”, “machine learning”, “determinis” and “sampl”. All seven return zero matches in both files, and zero again in GSA’s ACR and VPAT FAQ. What that proves is narrow, and it is enough to plan around: the instrument never names the technology, and neither the template nor the FAQ says how much of a product to exercise before answering a row. Provisions that reach output produced at runtime still exist. The vocabulary for pointing at it does not.

That is not a reason to leave rows blank. It is a reason to know which rows the panel touches, which of them your edition contains, and what a reviewer can replicate when the output changes on every run.

One label applies to everything below. No primary source consulted for this article states how to document a non-deterministic feature in an ACR. Where this article proposes a conformance level, that is ADA Compliance Pros’ proposed method, offered as analysis. It is not guidance from W3C, ITI, GSA or the Access Board. No published guidance from any of them on this question turned up in the research behind this article.

Four routes out of the panel, and what each one gets you

The first route is to exclude the panel and report the rest. Conformance Requirement 2 in WCAG 2.2, a W3C Recommendation of 12 December 2024, closes it: “Conformance (and conformance level) is for full web page(s) only, and cannot be achieved if part of a web page is excluded.” The VPAT prints the same rule above Table 1, scoping the criteria “for full pages, complete processes, and accessibility-supported ways of using technology as documented in the WCAG 2.0 Conformance Requirements.” A panel that opens on every page is part of every page.

The next two routes are both inside WCAG 2.2 section 5.4, which covers pages “that will later have additional content added” and offers two options. Read why the section exists before deciding whether it fits: “in these cases, it is not possible to know at the time of original posting what the uncontrolled content of the pages will be.” That is close to a vendor’s honest position about a model, and it is the strongest argument for reaching for 5.4 at all.

The second route is 5.4’s first option, which is monitoring rather than disclaiming. “A determination of conformance can be made based on best knowledge. If a page of this type is monitored and repaired (non-conforming content is removed or brought into conformance) within two business days, then a determination or claim of conformance can be made.” WCAG-EM 2.0 restates it as a question the evaluator has to answer: whether the content “is regularly monitored and repaired (within two business days), and whether non-conforming content is clearly identified as such in all the web pages in which it appears.” A moderated assistant with a triage queue is the one shape of AI panel that can argue this, and the arithmetic is unforgiving, because the clock runs against the page rather than against your release train. If the only fix for a class of bad output is a model swap or a prompt change on a monthly cadence, two business days is not met, and WCAG closes with the condition that decides it: “No conformance claim can be made if it is not possible to monitor or correct non-conforming content.”

The third route is 5.4’s second option, the statement of partial conformance, under which a page “does not conform, but would conform to WCAG 2.2 at level X if the following parts from uncontrolled sources were removed.” The first of its two conditions is the obstacle: “It is not content that is under the author’s control.” The reading in favor is real, since nobody can know at posting time which tokens the model will emit. Against it: you chose the model, wrote the system prompt, constrained the renderer and shipped the panel, and WCAG’s examples of uncontrolled content are things a site operator receives from other people, such as user comments, aggregated contributions and dynamically inserted advertisements. No source found in this research applies section 5.4 to model output either way. A reviewer who rejects the reading has the plain text of the condition on their side, and the second condition binds anyone who tries: the parts have to be “described in a way that users can identify.”

The fourth route is the one the template blesses. Under Best Practices for Authors, the VPAT says that “for complex products it may be helpful to separate answers into multiple reports,” and where you do, “it is required to explain this and how to reach the other reports in the Notes section of each report.” That works for an assistant sold as a separable surface. It does not rescue a panel embedded in pages the main report already covers.

One exception looks like relief and is not. If the panel ships inside desktop or mobile software, 36 CFR part 1194, Appendix A, E207.2 Exception 3 provides that “Non-Web software shall not be required to conform to Conformance Requirement 3 Complete Processes in WCAG 2.0.” That reaches multi-step processes. It leaves Conformance Requirement 2 untouched. The bridge document for that reader is WCAG2ICT, a W3C Group Note of 11 December 2025, whose section on SC 4.1.3 says the criterion “applies directly as written, and as described in Intent from Understanding Success Criterion 4.1.3.”

Four routes for handling an AI panel in a conformance report, and where each one fails. Excluding the panel claims the rest of the product while leaving the panel out, but WCAG Conformance Requirement 2 allows conformance only for full pages and cannot exclude a part, and a panel that opens on every page is part of every page. Monitoring under WCAG 2.2 section 5.4 allows a determination on best knowledge with bad output repaired, but requires non-conforming content to be fixed within two business days, which a monthly model swap or prompt change misses. A statement of partial conformance says the page would conform if the uncontrolled parts were removed, but requires parts that are not under the author's control and are described so users can identify them, and you chose the model and shipped the panel. A separate report gives the assistant its own document and requires the split and the route to the other reports to be explained in each report's Notes, and it does not rescue a panel embedded in pages the main report already covers.
Only the fourth route is written into the template, and it works only for an assistant sold as a separable surface.
View the data as a table
Exclude the panelMonitor and repair (5.4)Partial conformance (5.4)A separate report
What it claimsThe rest of the product, panel left outBest knowledge and bad output repairedIt would conform if the odd parts were cutThe assistant gets its own report
What the source demandsNothing. Full pages only, no part left outBad output fixed within two business daysParts you do not control, named for usersNotes in each report on where the other is
Where it failsA panel on every page is part of every pageA monthly model swap misses the clockYou chose the model and shipped the panelNot for a panel inside pages already covered

Which edition you file decides which questions exist

ITI states the version mapping on its own VPAT page: “WCAG 2.0 is incorporated into the 508 edition”, “WCAG 2.1 is incorporated into the EU edition”, “WCAG 2.2 is incorporated into the WCAG and INT editions”. Section 508 incorporates WCAG 2.0 by reference at 36 CFR part 1194, Appendix C, 702.10.1, naming the “W3C Recommendation, December 11, 2008”. That matters here because several criteria a generated surface stresses hardest arrived in WCAG 2.1. Counting exact row labels in the extracted text of both templates gives this:

Row label in the templateVPAT 2.5Rev 508 editionVPAT 2.5Rev INT edition
4.1.3 Status Messages01
1.4.11 Non-text Contrast01
2.5.3 Label in Name01
1.4.10 Reflow01
1.4.12 Text Spacing01
1.3.1 Info and Relationships11
1.4.3 Contrast (Minimum)11
2.2.1 Timing Adjustable11
2.2.2 Pause, Stop, Hide11
3.1.2 Language of Parts11
3.2.2 On Input11
4.1.2 Name, Role, Value11

A buyer who asks a 508-edition ACR what your assistant does about status messages is asking a document with no row for the question. Nothing is being hidden: the criterion is not part of the standard that report answers. Say so in the Notes rather than inventing a row. Only the 508 and INT templates were counted for the table above. That the WCAG and EU editions also carry the WCAG 2.1 rows follows from ITI’s version mapping, not from a row census. Which rule binds which buyer to which version is set out in our WCAG version by rule matrix.

One feature of the row itself is easy to miss and useful here. Inside a single WCAG row, the Conformance Level cell of the 508 edition is pre-split into four labeled lines, “Web:”, “Electronic Docs:”, “Software:” and “Authoring Tool:”. A panel that ships in a web app and in a desktop app can therefore carry two different verdicts in the same row without deviating from the template at all.

The progress line is a status message. The answer body is not.

Filing everything the panel announces under SC 4.1.3 Status Messages looks tidy, and W3C’s own explanation of the criterion rules it out.

The WCAG 2.2 glossary defines a status message as a “change in content that is not a change of context, and that provides information to the user on the success or results of an action, on the waiting state of an application, on the progress of a process, or on the existence of errors.” The Understanding document for SC 4.1.3, W3C supporting material rather than part of the Recommendation, draws the line: “the list of results obtained from a search are not considered a status update and thus are not covered by this success criterion. However, brief text messages displayed about the completion or status of the search, such as ‘Searching…’, ‘18 results returned’ or ‘No results returned’ would be status updates if they do not take focus or cause a page refresh.” Its list of things that are not status messages includes content added to the page in response to what the user supplied, which “do not meet the definition of status message.”

Split the row on that line. “Generating”, “Thinking”, “Rate limit reached” and “Response complete” are the waiting state and the progress of a process, so they are 4.1.3 material. The answer itself is content, and 4.1.3 does not reach it: the body is examined under 1.3.1 Info and Relationships, 1.4.3 Contrast (Minimum) and 4.1.2 Name, Role, Value, rows that both editions in the table above contain.

How the progress layer is announced is where a bare Supports stops being defensible. WAI-ARIA 1.2, a W3C Recommendation of 06 June 2023, defines the status role as a live region “not important enough to justify an alert”, carrying “an implicit aria-live value of polite”. For the running transcript, the log role fits closer: “a type of live region where new information is added in meaningful order and old information may disappear.” For a stream that should be announced once rather than token by token, aria-busy is the mechanism, and it is a permission rather than a guarantee: while it is true, assistive technologies “MAY ignore changes to content owned by that element and then process all changes made during the busy period as a single, atomic update when aria-busy becomes false.”

Then read the sentence that caps all of it. ARIA says politeness levels “serve as a strong suggestion to user agents or assistive technologies. The value may be overridden by user agents, assistive technologies, or the user.” A correct implementation cannot promise a particular announcement. Write what was implemented and what was observed, with browser and screen reader versions named, rather than a one-word verdict.

The self-rewriting response lands in criteria your report does not require

Start by ruling out the wrong row. SC 3.2.2 On Input is scoped to “Changing the setting of any user interface component”, and the federal ICT Testing Baseline for Web identifies the content for its On Input test as “All active form components.” A model revising its own output is neither.

The criteria that do describe the rewrite sit at Level AAA. SC 3.2.5 Change on Request requires that “changes of context are initiated only by user request or a mechanism is available to turn off such changes.” SC 2.2.4 Interruptions requires that “interruptions can be postponed or suppressed by the user, except interruptions involving an emergency.” Both appear only in Table 3 of the VPAT, the Level AAA table, where the 508 edition annotates every row to say the Revised Section 508 Standards do not apply. Section 508 requires Level A and AA, so the two criteria that name the problem most precisely sit outside the obligation the report documents. Table 3 is also the only place the template permits its fifth term, “Not Evaluated”, which “can only be used in WCAG Level AAA criteria.”

What is left inside Level A is 4.1.2 Name, Role, Value and the live-region defaults underneath it. ARIA’s default for aria-relevant is “additions text”, under which “text modifications and node additions are relevant, but that node removals are irrelevant”, and the default for aria-atomic is false, so “assistive technologies will only present the changed node to the user”. A response that replaces its own text under those defaults is re-announced in fragments, with the removal silent. Partially Supports, with a remark naming the behavior, is the honest answer.

SC 2.2.2 Pause, Stop, Hide is genuinely unsettled, and saying so beats picking a side. The criterion’s own text is conditional. It reaches “any auto-updating information that (1) starts automatically and (2) is presented in parallel with other content”, and it exempts updating that “is part of an activity where it is essential”. Baseline test 21.C is the federal test procedure built on that sentence, and it keeps all three conditions conjunctive: content is in scope only where it “Starts automatically, AND Is presented in parallel with other content, AND Is not part of an activity where it is essential”. A response stream starts when the user presses send, which is not obviously automatic; the Baseline’s gloss on “in parallel” describes a news flash beside a news video and text news articles, not the panel that is the thing being read; and the stream is arguably the activity itself, which brings the essential exception into play. No source resolves it. Put the reasoning in the remark.

Chips, names written at runtime, and two rows people forget

A suggested-action chip is a control whose accessible name exists only after the model has produced it. ARIA’s button role carries “Accessible Name Required: True”, with the name taken from the element’s own contents or from the author. The federal Baseline test for control names adds the instruction that matters here: “If the name of the user control changes on user interaction with the web content or application, repeat the previous test steps and check that the accessible name is correct after the change.” Its companion test says the same for state. Supports is defensible when the chip’s accessible name comes from its own rendered text by a path that does not vary. Where the name is composed separately from the visible label, Partially Supports is accurate, and on an edition built on WCAG 2.1 or 2.2 that mismatch has its own row at 2.5.3 Label in Name: “the name contains the text that is presented visually”.

Two more rows get missed. SC 3.1.2 Language of Parts requires that “the human language of each passage or phrase in the content can be programmatically determined”. An assistant that answers a Spanish question in Spanish inside an English page, with no lang attribute on the generated passage, does not meet it. And if the assistant can be driven by voice, the governing provision is not a WCAG row at all: 36 CFR part 1194, Appendix C, 302.6 requires “at least one mode of operation that does not require user speech.” That is a functional performance criterion, answered in a different table, and we have written separately about the functional performance criteria rows.

Four AI panel behaviors mapped to the criterion they touch, whether the 508 edition has that row, and the level ADA Compliance Pros proposes. The progress and status line is SC 4.1.3 Status Messages, the waiting state and the progress of a process; the 508 edition has no row for it at all; the proposed answer is not a bare Supports but a note naming what was implemented and observed. The self-rewriting answer falls to SC 4.1.2 Name, Role, Value, since the criteria that name the rewrite best sit at Level AAA; 4.1.2 is a row in both editions while the AAA criteria appear only in Table 3, where Section 508 does not apply; the proposed level is Partially Supports with a full remark. The suggested-action chip falls to SC 4.1.2 and to SC 2.5.3 Label in Name; 4.1.2 is in both editions and 2.5.3 only on an edition built on WCAG 2.1 or 2.2; the proposed level is Supports where the accessible name is the chip's own visible text. An answer in another language falls to SC 3.1.2 Language of Parts, a row in both editions counted, and the proposed level is Does Not Support where the generated passage carries no lang attribute.
The verdict row is ADA Compliance Pros’ proposed method, offered as analysis; no primary source consulted states how to document a non-deterministic feature.
View the data as a table
Progress and status lineSelf-rewriting answerSuggested-action chipAnswer in another language
Criterion it touches4.1.3 Status Messages: waiting and progress4.1.2 Name, Role, Value. The rest is AAA4.1.2, and 2.5.3 Label in Name3.1.2 Language of Parts
Row in the 508 edition?No row at all in that edition4.1.2 yes. The AAA rows are Table 3 only4.1.2 yes. 2.5.3 only on 2.1 and 2.2Yes, a row in both editions counted
Level ADACP proposesNot a bare Supports. Name what you observedPartially Supports, plus a full remarkSupports if the name is its own visible textDoes Not Support with no lang attribute

Evidence a reviewer can replicate

The VPAT requires a description of “evaluation methods used to complete the VPAT for the product under test”, and its best practices add that “if a published test method was used, provide name, publisher, URL link of the test method.” Two are worth naming here, each with a caveat about what it is.

The first is federal. The ICT Testing Baseline portfolio gives Baseline for Web version 3.1, published April 1, 2024, and its test procedures are the ones quoted above. Read what the portfolio says about itself before citing it as coverage. It is “a comprehensive set of test components that a Section 508 conformance test process should include”, and it is expressly not “a step-by-step testing procedure or methodology”. More to the point for an AI panel: “the Baselines are mapped only to Section 508 (and WCAG 2.0) requirements.” The full text of the web baseline tests contains zero occurrences of “status message” and zero of “4.1.3”, and no test for 2.5.3 Label in Name or 1.4.11 Non-text Contrast either. Naming the Baseline in Evaluation Methods Used is accurate. It is also not coverage of the WCAG 2.1 criteria this article says a generated surface stresses hardest. If you are filing an INT or WCAG edition, say what you used for those rows.

The second is WCAG-EM 2.0, a W3C Group Note of 23 July 2026, not a standard and not something Section 508 incorporates. It is still the only published methodology found in this research that speaks to a surface like yours. Its sampling step requires “samples that reflect all identified (1) common views, (2) essential functionality, (3) types of samples, (4) technologies relied upon, and (5) other relevant samples”, and its list of sample types names “dynamic content” and content that changes “depending on the user, device, browser, context, and settings”. Its random sample rule carries the only sample-size number in any source consulted here: “the number of samples to randomly select is 10% of the structured sample set selected through the previous steps.”

Two of its steps decide whether a conversation is testable, and they are not interchangeable. Step 3.3, on including complete processes, carries the note that “in most cases the web address (URL) will not be sufficient to identify the sample in a complete process”, and asks for the actions needed to move from one sample to the next. Step 5.2 is where the prompt log lives, and WCAG-EM marks that step optional in its own requirement line: “Archive the samples evaluated, and record the evaluation tools, web browsers, assistive technologies, other software, and methods used to evaluate them (optional).” Its bullets ask for a “Description of the settings, input, and actions used to generate or navigate to the samples” and the “Names and versions of the evaluation tools, web browsers and add-ons, assistive technology, and other software used”. Optional in WCAG-EM is not optional in your report, because the VPAT requires an Evaluation Methods Used section either way. For an AI panel that archive is a prompt log: the prompt text, the account and settings state, the path taken to reach the panel, the tool, browser and assistive technology versions, and the transcripts.

Now the gap. No source found in this research states how many times a generated response must be exercised before a verdict is defensible. WCAG-EM sizes samples by breadth of views, not by repeat runs of one view. Any number of runs in your evaluation methods section is your own judgment, and it should read as a stated protocol rather than as conformity with a published rule.

The two test methods worth naming for an AI panel, compared. The ICT Testing Baseline for Web is version 3.1, published April 1, 2024, a comprehensive set of test components a Section 508 conformance test process should include, expressly not a step-by-step testing procedure or methodology; its baselines are mapped only to Section 508 and WCAG 2.0 requirements; the full text of the web baseline tests contains zero occurrences of status message and zero of 4.1.3, and no test for 2.5.3 Label in Name or 1.4.11 Non-text Contrast; naming it is accurate but it is not coverage of the WCAG 2.1 criteria. WCAG-EM 2.0 is a W3C Group Note of 23 July 2026, not a standard and not something Section 508 incorporates; its sampling step requires samples reflecting common views, essential functionality, types of samples, technologies relied upon and other relevant samples, and its random sample rule sets 10 percent of the structured sample set; it sizes samples by breadth of views, not by repeat runs of one view, so no source states how many times a generated response must be exercised; its Step 5.2 archive is marked optional in WCAG-EM but the VPAT requires an Evaluation Methods Used section either way.
Name either one in Evaluation Methods Used and the caveat travels with it, which is why the rows a generated surface stresses need their own sentence.
View the data as a table
ICT Testing Baseline for WebWCAG-EM 2.0
What it isFederal test components, Baseline for Web 3.1, published April 1, 2024A W3C Group Note of 23 July 2026, not a standard, not incorporated by 508
What it reachesMapped only to Section 508 and WCAG 2.0 requirementsSampling that reflects common views, essential functionality and dynamic content
Where it runs outZero tests for status messages, 4.1.3, 2.5.3 or 1.4.11 Non-text ContrastSizes samples by breadth of views, never by repeat runs of one view
How to cite it honestlyAccurate to name, but it is not coverage of the WCAG 2.1 rowsIts Step 5.2 archive is optional there and required by your report anyway

What the remark has to contain

The template is specific, and the specification is the same for an AI panel as for anything else. When the conformance level is “Partially Supports” or “Does Not Support”, it says the remarks should identify:

  • “The functions or features with issues”
  • “How they do not fully support”
  • “If the criterion does not apply, explain why.”
  • “If an accessible alternative is used, describe it.”

Best practices add application dependencies such as operating system and browsers relied on, known workarounds, and how the customer can find more information.

For a non-deterministic feature, that becomes a remark naming the panel and its boundary, stating the shortfall in terms a tester can reproduce, naming the browser and assistive technology combinations the claim was verified against, giving the workaround, and pointing at a tracked defect. It is a long cell, and it is what stands between your report and a reviewer’s conclusion that the row was guessed.

A row that says Partially Supports is also not automatically a lost bid, which is the question behind the questionnaire. Where ICT conforming to the standards “is not commercially available”, 36 CFR part 1194, Appendix A, E202.7 directs the agency to “procure the ICT that best meets the Revised 508 Standards consistent with the agency’s business needs”, and E202.7.1 makes the responsible official document in writing “which provisions cannot be met” and the basis for that determination. A remark precise enough to be quoted into that written determination is worth more to the buyer than a Supports they cannot defend.

Two terms repay precision. ITI defines Supports as functionality that “has at least one method that meets the criterion without known defects or meets with equivalent facilitation”, and Partially Supports as “some functionality of the product does not meet the criterion.” Deviate from those definitions and the template requires you to say so in the Notes. GSA’s FAQ adds the sentence that turns a model change into a reporting event: “Every time your product is changed or updated (e.g. version change, bug fix, etc.), an updated ACR may be required to address any changes in the product’s accessibility.” Our guide to what makes an ACR credible covers the document around these rows.

The remark cell for a non-deterministic feature, on a Partially Supports row, broken into five parts. Panel and boundary: which feature the row answers for. The shortfall: how it falls short, in a tester's terms. Verified pairs: browser and assistive technology. The workaround: plus the dependencies relied on. A tracked defect: the defect record the row points at.
It is a long cell, and it is what stands between the row and a reviewer’s conclusion that the verdict was guessed.
View the data as a list

The remark cell: For a non-deterministic feature, on a Partially Supports row

  • Panel and boundary: Which feature the row answers for
  • The shortfall: How it falls short, in a tester’s terms
  • Verified pairs: Browser and assistive technology
  • The workaround: Plus the dependencies relied on
  • A tracked defect: The defect record the row points at

Three questions nobody has answered

State these in the Notes rather than papering over them.

Does swapping the model trigger a new ACR? GSA says a product change may require an updated report, and gives version changes and bug fixes as its examples. Nothing found says whether replacing the underlying model, editing a system prompt or changing a decoding parameter counts. Write down the answer you chose and the date you chose it, and use the mechanism the template already gives you: “If a report is revised, change the report date and explain the revision in the Notes section. Alternately, create a new report and explain in the Notes section that it supersedes an earlier version of the report.”

Has any federal agency ruled on how an ACR should document an AI feature? Searches of Section508.gov and the Federal Register for this article surfaced no rulemaking, no guidance document and no enforcement action addressing an ACR row for an AI feature. The one federal statement found that joins generative AI to Section 508 methodology is a recommendation in the FY 2025 Governmentwide Section 508 Assessment, and its subject is agencies using AI to produce documents and web pages. To keep that content conformant before publication, it says, “it is beneficial if agency testing methodologies are aligned with the Baselines for evaluation.” That is not about a vendor documenting an AI feature. The same page does tell buyers to evaluate ICT “through accessibility conformance reports (ACRs) to validate the accuracy of vendor conformance claims”, which is the pressure your report is about to meet.

One live federal item belongs here even though it settles nothing. On 24 June 2026 GSA published Information Collection; Accessibility Conformance Report (ACR) Repository, Federal Register document 2026-12667, a notice and request for comments announcing that under the Paperwork Reduction Act it “will be submitting to the Office of Management and Budget (OMB) a request to review and approve a new information collection requirement regarding the Accessibility Conformance Report (ACR) Repository.” Comments closed on 24 August 2026. It is an information collection request rather than a rule, and nothing in it addresses AI features. Watch it anyway: if OMB approves the collection, an ACR becomes something a federal repository holds rather than an attachment to one bid.

Is there a requirements document for conversational interfaces? There is a W3C document, and its status matters. Natural Language Interface Accessibility User Requirements is a Group Draft Note of 03 September 2022, and its abstract says of itself: “This document is not a collection of baseline requirements.” Cite it for the user needs it names, such as the requirement that a user “can review the entire history of the conversation”, never as a conformance target.

Where this stops

Write the rows, write the remarks, name the method, date the report. What none of that settles is whether a given report creates contractual or legal exposure for your company, which is a question for your counsel and not for a testing firm.

If you want the rows written against the panel you actually shipped, with the evaluation methods section drafted so a federal reviewer can reproduce the result, that is what our VPAT and ACR work does, and it is the part of our SaaS and software provider practice procurement reads first. Send the panel, the edition your buyer asked for, and the release process behind the model. The first thing back is a list of the rows where a verdict is standing in for a remark.