The accessibility audit report a buyer should refuse to accept
Two audits, one price, two different deliverables
Take a worked example. Two firms quote the same website at the same number, and both proposals say “WCAG 2.2 Level AA audit, full report.” One delivers a document a developer can work from on Monday: every view in scope named, the sample listed with the method used to pick it, each failure tied to a numbered success criterion, the steps that reproduce it. The other delivers a spreadsheet of rule names with a severity column, no page addresses, and no way to reproduce anything.
Both firms can defend that deliverable against the words in the contract, because the contract never said which document was being bought. The gap surfaces at acceptance.
This page is about the artifact, not the process and not the price, which is covered in what an accessibility audit costs. What follows splits accessibility audit deliverables three ways: the fields the published methodology says a report must carry, the fields it marks optional, and the fields no methodology requires at all, which will not arrive unless you write them into the statement of work.
The methodology everybody cites is a Group Note
A proposal can cite “WCAG-EM” as though it were a rule with teeth. It is not one.
WCAG Evaluation Methodology 2.0 was published on 23 July 2026. Its Status of This Document section says it “was published by the Accessibility Guidelines Working Group as a Group Note using the Note track,” and the next sentence removes the ambiguity: “This Group Note is endorsed by the Accessibility Guidelines Working Group, but is not endorsed by W3C itself nor its Members.” The introduction adds that it “does not in any way add to or change the requirements defined by the normative WCAG 2 standard.” Its predecessor, WCAG-EM 1.0 of 10 July 2014, is still on the W3C site as a Working Group Note, so a report naming no version has told you nothing.
That does not make the methodology useless to a buyer. It makes it the wrong kind of leverage: you cannot accuse a vendor of violating a standard by omitting a field, but you can put the field in the contract, and the methodology gives you defensible language for doing it.
One of the five reporting requirements is not optional
Step 5, the reporting step, carries five numbered Methodology Requirements. Four of them, 5.2 through 5.5, are labeled “(optional)” in their own headings. Only Methodology Requirement 5.1 is not:
Document each outcome of the steps defined in Step 1: Define the evaluation scope, Step 2: Explore the target digital product, Step 3: Select a representative sample set, and Step 4: Evaluate the selected sample set.
That pulls the whole evaluation into the report by reference, for a reason worth quoting in a negotiation: documenting every previous step “is essential” for “transparency, replicability of the evaluation results and justifications for any statements made.”
Underneath it, Step 5.1 prints an “Include at least the following” list, and some entries inside that list carry an “Optional:” prefix of their own. Here is the whole list, with each entry marked as the methodology marks it. This is the list to paste into acceptance criteria.
| Field | Source step | Marked |
|---|---|---|
| Name of the evaluator, optionally including contact details | Step 5.1 | Owed |
| Name of the evaluation commissioner | Step 5.1 | Owed |
| Date of the evaluation, as completion date or duration period | Step 5.1 | Owed |
| Version number or unique identifier of the evaluation | Step 5.1 | Optional |
| List of dates, such as the initial report and repeat evaluations | Step 5.1 | Optional |
| Name of the person, team or organization responsible for the product | Step 5.1 | Optional |
| Methodology used for evaluation | Step 5.1 | Optional |
| Scope of the digital product | Step 1.1 | Owed |
| Conformance target | Step 1.2 | Owed |
| Accessibility support baseline | Step 1.3 | Owed |
| Additional evaluation requirements, where any were defined | Step 1.4 | Owed |
| Technologies relied upon | Step 2.4 | Owed |
| Common views of the digital product | Step 2.1 | Optional |
| Essential functionality of the digital product | Step 2.2 | Optional |
| Variety of sample types | Step 2.3 | Optional |
| Other relevant samples | Step 2.5 | Optional |
| Pages or views selected through structured sampling | Step 3.1 | Owed |
| Randomly selected samples and selection method used | Step 3.2 | Owed |
| Complete processes selected | Step 3.3 | Owed |
| Evaluation outcomes for all initial samples | Step 4.1 | Owed |
| Evaluation outcomes for all complete processes | Step 4.2 | Owed |
| Comparison of the structured and random sample sets | Step 4.3 | Owed |

View the data as a table
| Owed by default | Marked optional in the source | A contract term only | |
|---|---|---|---|
| Example fields | Scope, conformance target, accessibility support baseline, technologies relied upon, the samples and the outcomes | Version number, methodology used, common views, essential functionality, variety of sample types | Steps to reproduce, severity, every failure occurrence, remediation owner, effort estimate, due date |
| How the source marks it | Listed in Step 5.1 with no prefix; MR 5.1 requires every Step 1 to 4 outcome to be documented | Carries an “Optional:” prefix of its own inside the same Step 5.1 list | A non-normative note, a “may request” sentence, or nothing at all |
| What arrives if you say nothing | The field, and grounds to send the report back without it | Nothing; the vendor can omit it and still follow the methodology | Nothing; a vendor who omits all of it has not broken the methodology |
| Where to write it | Acceptance criteria, pasted from the Step 5.1 field list | Named individually in the statement of work | The statement of work, under Methodology Requirement 1.4 |
Three fields decide whether a report is usable on delivery, and not one of them appears in that list.
Steps to reproduce and severity sit in a note, not a requirement. They appear in a non-normative note under Step 5.1, and the verb is not “must”: “clear issue descriptions, steps to reproduce, severity of the findings, screenshots and/or videos can help teams resolve issues more quickly.” A vendor who omits both has not broken the methodology.
Every failure occurrence, rather than one example, is something you request. Step 5.1 says “an evaluation commissioner may request a report indicating every failure occurrence for every sample, more information about the nature and the causes of the identified failures, or repair suggestions to remedy the failures.” The hook is Methodology Requirement 1.4, and Step 1.4 adds that such requirements “need to be clarified early on and documented. This also needs to be reflected in the resulting report.” Raising it at delivery is too late.
A remediation owner is not a field in any methodology. No owner, effort estimate or due date appears in the report fields of WCAG-EM 2.0, the ICT Testing Baseline, DHS Trusted Tester v5.1.3, GSA’s test report guidance, 36 CFR part 1194, 28 CFR part 35 or FAR subpart 39.2. Priorities and deadlines do exist in enforcement agreements, which bind the entity that signed them rather than your vendor, and that distinction is worked through at the end of this page. If your ticket workflow needs those columns, they are yours to specify.
Scope, level and baseline: the three lines that make the rest mean anything
Step 1 has three non-optional sub-requirements, and MR 5.1 drags all three into the report.
Methodology Requirement 1.1 asks the evaluator to define the product “so that for each view it is unambiguous whether it is within the scope of evaluation or not.” Hold the delivered scope line against that one test. “The marketing site” fails it, because two people reading it will draw the boundary in two places and neither can be shown wrong.
Methodology Requirement 1.2 asks only to “Select a target WCAG 2 conformance level (A, AA, or AAA) for the evaluation.” Read it twice for what it does not ask. It wants a level, not a version, so a report can satisfy MR 1.2 without ever naming which WCAG it tested against. The version surfaces only in the optional evaluation statement at Step 5.3. Name it yourself, matched to the rule your organization is measured against.
Methodology Requirement 1.3 is the accessibility support baseline: “Define the web browser, assistive technologies and other user agents for which features provided on the digital product are to be accessibility supported.” That is a product-level statement, not a per-finding one. Per-finding environment records sit in the optional MR 5.2, of which Step 5.2 says: “This recording is typically kept internal and not shared by the evaluator unless otherwise agreed on in Step 1.4.” A browser and screen reader pairing on every individual finding is a contract term, not a default.
Reading the sample section of a delivered report
A report has to make its own coverage legible, and there are only two shapes it can take. Either the evaluator sampled, in which case the Step 5.1 field list makes the structured samples, the randomly selected samples with “selection method used,” and the complete processes all reportable. Or the evaluator skipped sampling and covered the whole product, which the methodology permits where that is feasible, and then the report has to say so. A document that does neither has left you unable to say what was looked at.
Where sampling was used, one number is checkable on the face of the report. Step 3.2 puts the random set at “10% of the structured sample set selected through the previous steps,” added on top rather than carved out of it. Whether the structured set underneath it was the right size is a question you settle before the work starts, in the statement of work, not in the margin of a delivered document.
One further test applies to process findings. Methodology Requirement 3.3 covers processes: “Include all samples that are part of a complete process in the selected sample set.” Its note is what to quote when a report hands you bare URLs for a checkout flow:
In most cases, it is necessary to record and specify the actions needed to proceed from one sample to the next in a sequence to complete a process so that they can be replicated later. … In most cases the web address (URL) will not be sufficient to identify the sample in a complete process.
What a single finding has to carry
There is a real, published, sworn finding to measure against, which beats a constructed example. In February 2026 the United States filed a Statement of Interest in Alcazar v. Fashion Nova, Inc., No. 4:20-cv-01434-JST (N.D. Cal.), attaching a declaration from an expert who had reviewed the settlement claims website. She declared her tools once, up front: “To conduct my review, I used five testing tools: NonVisual Desktop Access, JAWS 2024 Screen Reader, iOS Voice Over, Axe Accessibility Tool, and Color Contrast Analyser.”
One of her findings reads in full:
Redundant “Submit” Button That Did Not Work. Third, when navigating the claims form using a screen reader, a blind user heard two “Submit” buttons. A sighted user of the claims form only saw one button. Only one of the “Submit” buttons announced by a screen reader actually submitted the form; the other did not. See WCAG 2.1, 1.3.1, Info and Relationships. Therefore, a blind claimant could have selected the wrong button, and incorrectly assumed their form was submitted.
Four things do the work there: the interaction that triggers the failure, the observed behavior set against the sighted behavior, the criterion named and numbered, and the consequence for a user trying to finish a task. Another finding carries a measured value rather than an adjective, reporting a focus indicator whose contrast “was too low,” then: “it had a ratio of 1.4:1. … Best practice would be a color contrast ratio of 3:1 against the white background. See WCAG 2.1, 1.4.11, Non-text Contrast.”
Now the careful part. That declaration pairs no browser or operating system with any individual finding, assigns no severity rating, and never says who should fix anything. That is worth knowing, and it is worth less than it looks: a declaration is written to persuade a judge, and the fields it leaves out are the fields a judge does not need. Set against GSA’s own report guidance, further down this page, which asks a tester to record the test environment, how to replicate each defect and “the criticality or severity of each defect,” the honest conclusion is narrower than “nobody does this.” No methodology owes you these fields, one federal guidance document asks for them, and practice is split. So buy them.

View the data as a table
| What a thin report prints | What the filed declaration carried |
|---|---|
| ”Buttons not labeled correctly. Severity: Medium. WCAG: 1.3.1.” | The interaction that triggers the failure |
| A severity value from no stated rubric, added rather than found | The observed behavior set against the sighted behavior |
| No view named, and no path to reach that state | The criterion named and numbered: WCAG 2.1, 1.3.1, Info and Relationships |
| No observed behavior, and no contrast with sighted behavior | The consequence for a user trying to finish a task |
| No user consequence | A measured value rather than an adjective: a ratio of 1.4:1 |
| No route to reproduction | The five testing tools declared once, up front |
Strip that finding to what a thin report prints and you get “Buttons not labeled correctly. Severity: Medium. WCAG: 1.3.1.” Six things went with the stripping: the view, the path to reach that state, the observed behavior, the contrast with sighted behavior, the user consequence, and any route to reproduction. One thing was added, a severity value from no stated rubric.
Severity has two sources, and neither one is WCAG
The severity column carries the least behind it of anything in a report.
WCAG 2.2 has no rating scheme, and Step 5.4 explains why: “there is currently no single metric that is known to address the required reliability, accuracy, and practicality. … For this and other reasons WCAG 2 does not provide a rating scheme.”
Two four-level scales fill the gap, from different places. axe-core’s API documentation describes an impact property that “Can be one of ‘minor’, ‘moderate’, ‘serious’, or ‘critical’ if the Rule failed or null if the check passed.” GSA’s test report guidance names a different four: it asks a tester to explain “the criticality or severity of each defect. Common levels include: Critical, High, Moderate or Medium, and Low.” Two scales, four levels each, different labels, no published bridge between them and no published rubric for either. That is why the rubric has to travel with the score.

View the data as a table
| WCAG 2.2 | axe-core impact | GSA report guidance | |
|---|---|---|---|
| Levels offered | None; no rating scheme at all | Four | Four |
| Labels | None | minor, moderate, serious, critical | Critical, High, Moderate or Medium, and Low |
| Where it comes from | The WCAG 2 standard itself | axe-core’s API documentation | GSA’s test report guidance for federal testers |
| Published rubric | No scheme, so no rubric | None published | None published |
US regulation contains one severity-shaped test, and it is not a vendor’s to use. 28 CFR 35.205 deems a public entity compliant where noncompliance “has such a minimal impact on access that it would not affect the ability of individuals with disabilities” to use its web content or mobile app, “in a manner that provides substantially equivalent timeliness, privacy, independence, and ease of use,” to access the same information, engage in the same interactions, conduct the same transactions, and otherwise participate in or benefit from the same services, programs and activities. It measures a public entity’s own posture, not a finding’s rank, and nobody has published a bridge between the two.
If you want severity, buy the rubric with it. Step 5.4 supports that: “Whenever a score is provided, it is essential that the scoring approach is documented and made available to the evaluation commissioner along with the report, to facilitate transparency and repeatability.”
Federal buyers have more to hold a report to
Federal work has mandatory report fields, and they live in a test process document rather than in a regulation.
The ICT Testing Baseline for Web sets “the minimum tests and evaluation guidance that determine whether Web content meets Section 508 requirements” and says it “is not intended to be a test process itself.” Its acceptance sentence is the useful one: “all test processes that claim to align to this baseline must include all baseline tests and provide baseline test results,” and “Agency-specific non-baseline tests must be identified, and these results must be reported separately from the baseline results.” One blended results table fails that.
The DHS Trusted Tester process is a manual approach that section508.gov describes as aligning with the ICT Testing Baseline, and it adds that agencies adopting it “only accept test results from individuals who have been certified as Trusted Testers.” Its current edition, version 5.1.3 of April 2024, states five duties WCAG-EM 2.0 does not impose:
- “There are 63 Test Conditions for evaluation in this test process. Each Test Condition must have a test result for testing to be considered complete.” An older 5.0 edition still online says 66, and carries a banner telling readers to use 5.1 or above.
- The deliverable has a floor format: “Trusted Tester results must be provided at minimum following the Accessibility Conformance Report format from the IT Industry consortium. However, the ACR format must be supplemented with specific Trusted Tester test outcomes.”
- Every Accessibility Conformance Report “must provide: Clear identification of the test process used to return conformance results. Clear indication of the scope of testing. Clear documentation of the test environment(s).”
- Outcomes come from a stated set: “In general, the possible outcomes for Test Conditions and web requirements are PASS, FAIL, DOES NOT APPLY, or NOT TESTED. Any results of FAIL should also include clear information identifying the location of the failure and, when feasible, clear information illustrating the content or information that resulted in the FAIL result.”
- Silence is not permitted: “If a tester cannot complete a test, a note should be added to the test report indicating ‘This test could not be performed’, with a detailed explanation of the issue.”
That last duty, the one about silence, has no equivalent in WCAG-EM 2.0 and is worth copying into a private-sector statement of work word for word. It turns an untested area from an invisible gap into a line item.
The second duty is the one to read twice, because it names a document. An ACR is not a test report, and section508.gov draws the line in one sentence: “An Accessibility Conformance Report (ACR) provides an overview of a product’s conformance to Section 508. In contrast, a Section 508 test report offers a more detailed, developer-oriented document to assist product teams in enhancing Section 508 conformance.” A buyer who asks for an ACR and stops there has asked for the summary and not the working document.
Two cautions. Baseline alignment is self-asserted for now: GSA’s Baseline Alignment Framework says its working group “is still developing test cases and guidance for test process and tool owners to evaluate alignment to the ICT Testing Baseline,” and “More detailed guidance is still to come.” And one baseline test cannot fail: 24.A-Parsing carries the instruction “No testing necessary” and the result “Baseline Test 24.A-Parsing passes.”
The federal field list for a report that is not an ACR
GSA publishes a page called “Essential Elements of an Accessibility Test Report”, reviewed in July 2025, and it is the closest thing in federal material to a field list for a report that is not an ACR. It binds nobody. It opens: “However you create an accessibility test report, there are essential elements every report should contain.” Four groups follow, each introduced with “At a minimum, include.”
- Product information. Product name, the specific version, and a brief description, because “specific product information must be included.”
- Tester information. Tester names, the organization or vendor name, contact details, and “Tester credentials, if any, to denote subject matter expertise in the area of testing, such as Trusted Tester ID.”
- Report overview details. Report date, date of evaluation, report version, then the method: “Specify what tools and test methodologies were used to complete testing,” “Specify the operating system, browser product or version used,” and “Specify the test scope, including what was tested, how many pages, what may have been omitted in test scope.”
- Test results. A conformance outcome for every applicable standard, and for the ones that do not apply too: “A completed test report should have a conformance outcome listed for each standard even if the standard does not apply.” Per defect, enough to act on: “what the defect is, when or where the defect occurs, and the criticality or severity of each defect,” a screenshot, a code snippet where one applies, and “sufficient detail to explain how developers can remediate the defect.”
Read that against the methodology’s own list and the difference is the argument of this page. Severity, replication steps, test environment, screenshots, code snippets and remediation guidance are asked for by a federal agency writing for its own testers, and are optional, buried in a note, or absent in the document a commercial proposal cites. Neither one binds a private-sector vendor. Both give you sentences to paste.
GSA’s formatting section adds one more line worth a clause of its own, “Machine-readable formats facilitate comparison capabilities and tool integration,” which is the sentence to cite when findings have to land in a tracker without being retyped.
Four conditions for sending the report back
One claim the methodology forecloses outright, in the section on its relation to conformance claims: WCAG 2 conformance claims “cannot be made for entire websites based upon the evaluation of a selected sub-set of web pages and functionality alone, as it is always possible that there will be unidentified conformance errors on these websites. … Thus, in the majority of situations, using this methodology alone does not result in being able to make WCAG 2 conformance claims.” WCAG 2.2 is equally unforgiving a level down: “Conformance (and conformance level) is for full web page(s) only, and cannot be achieved if part of a web page is excluded.” A cover page reading “Your site is WCAG 2.2 AA conformant” is a claim the underlying work cannot support.

View the data as a list
- Scope, conformance level or baseline missing: The three non-optional parts of Step 1, and MR 5.1 requires the outcome of every Step 1 sub-step to be documented.
- No random sample, and no statement of full coverage: Step 3.2 puts the random set at ten percent of the structured set, added on top. A report that skipped sampling has to say so.
- A criterion not met with no example attached: Step 5.1 asks for at least one example for each conformance requirement and WCAG 2 Success Criterion not met.
- Whole-product conformance claimed from a sample: Conformance claims cannot be made for entire websites from the evaluation of a selected sub-set of pages and functionality alone.
- Federal only: a Test Condition with no result: Or a test skipped without the explicit “This test could not be performed” note and its explanation.
Each of the four is refusable on the face of the document, with no retesting required.
- Scope, conformance level or accessibility support baseline is missing or ambiguous. These are the three non-optional parts of Step 1, and MR 5.1 requires the outcome of every Step 1 sub-step to be documented.
- There is no random sample set, and no statement that the entire product was evaluated instead. The MR 5.1 field list calls for the samples and the selection method used, and a report that skipped sampling has to say that is what happened.
- A conformance requirement or success criterion is recorded as not met with no example attached. Step 5.1: “Reports should include at least one example for each conformance requirement and WCAG 2 Success Criterion not met.”
- The report asserts WCAG conformance for the whole product on the strength of a sample.
A fifth applies to federal work only: reject a Section 508 report that leaves any Test Condition without a result, or that skips a test without the explicit “This test could not be performed” note and its explanation.
What is not settled
No US rule prescribes the contents of a third-party audit report. 28 CFR 35.200 requires public entities to comply with WCAG 2.1 Level A and AA from April 26, 2027, or April 26, 2028 for smaller entities and special district governments, but it requires an outcome, not a document. Section 508 requires conformance to WCAG 2.0 Level A and AA for electronic content at provision E205.4, and for software at E207.2, of 36 CFR part 1194 Appendix A, again with no document specified. The report fields that bind a deliverable are in Trusted Tester v5.1.3, and only for agencies that adopted that process. GSA’s essential elements page is guidance and says so in its own voice, describing what a report “should contain.”
No court opinion was found construing what a third-party audit report must contain. Enforcement instruments are a different matter. At least one of them writes the report contents down. The 2017 resolution agreement between the Department of Education’s Office for Civil Rights and Santa Clara University, OCR Reference No. 09-17-2106, provides: “Within 90 calendar days of the most recent Audit, the Recipient will submit to OCR documentation of the steps taken by the Auditor during the Audit, a description of the outreach it undertook and the input it received, and a detailed accounting of the results of the Audit.” The next section requires a Corrective Action Plan that “will set out a detailed schedule for addressing problems, taking into account identified priorities, with all corrective actions to be completed within 18 months of the date OCR approved the Corrective Action Plan.” Auditor method, outreach, results, priorities and a deadline, written as an obligation. It binds the university that signed it and nobody else, and it is the closest thing in US practice to an enforcement-tested account of what an audit has to produce. The Alcazar declaration, by contrast, is evidence of what one expert wrote, not a holding about what a report owes.
For a private-sector buyer there is no federal standard to measure a report against at all. Footnote 4 of that same filing says: “The United States does not endorse WCAG as the appropriate or necessary standard for the provision of auxiliary aids and services under Title III of the ADA.” Whatever the contract names is the standard.
And no published dataset counts how many delivered reports omit the fields above. The owed-versus-optional split is documented here from the sources; how the market behaves against it is not measured, and no number should be attached to it.
One boundary. Deciding whether a delivered report is contractually conforming, and what to do if it is not, is work for your counsel and your contracting officer. What sits on this side of the line is which fields exist, which are owed by default, and which have to be written into the statement of work before testing starts. That is what our WCAG audits and testing engagements are scoped around.
Questions buyers ask before signing an audit contract
Which accessibility audit deliverables does WCAG-EM 2.0 say you are owed?
The outcome of every step, reported. Methodology Requirement 5.1 is the only non-optional reporting requirement, and it requires the outcomes of Steps 1 through 4 to be documented. Step 5.1 turns that into a field list: evaluator and commissioner names, the date of the evaluation, the scope, the conformance target, the accessibility support baseline, any additional requirements agreed at Step 1.4, the technologies relied upon, the structured sample, the random sample with the method used to pick it, the complete processes, and the evaluation outcomes from Steps 4.1, 4.2 and 4.3. Entries in the same list marked “Optional:” include the methodology used, the common views, the essential functionality and the sample types.
Does a WCAG-EM audit report have to list the browsers and screen readers used?
At the product level, yes. Methodology Requirement 1.3 requires an accessibility support baseline naming the browsers, assistive technologies and other user agents the product is meant to support, and MR 5.1 requires that outcome to be reported. Per-finding environment records are different: they sit in the optional MR 5.2, and Step 5.2 says that record “is typically kept internal and not shared by the evaluator unless otherwise agreed on in Step 1.4.” GSA’s guidance for federal test reports does ask a tester to “specify the operating system, browser product or version used,” so the field has a federal home even though the methodology leaves it optional.
Can an accessibility audit report claim the whole website conforms to WCAG?
Not from a sample. WCAG-EM 2.0 states that conformance claims “cannot be made for entire websites based upon the evaluation of a selected sub-set of web pages and functionality alone,” and that “in the majority of situations, using this methodology alone does not result in being able to make WCAG 2 conformance claims.”
Does WCAG define a severity scale for accessibility audit findings?
No. WCAG-EM 2.0 states that “For this and other reasons WCAG 2 does not provide a rating scheme.” Two four-level scales circulate instead, from outside WCAG: axe-core’s impact property, with “minor”, “moderate”, “serious” and “critical”, and GSA’s federal report guidance, which names “Critical, High, Moderate or Medium, and Low.” If a report carries severity, ask for the rubric: Step 5.4 says any scoring approach should be “documented and made available to the evaluation commissioner along with the report.”
Who is responsible for fixing each issue in an accessibility audit report?
The report does not say, unless you paid for it to. No remediation owner, effort estimate or due date field appears in the report fields of WCAG-EM 2.0, the ICT Testing Baseline, Trusted Tester v5.1.3, GSA’s essential elements guidance, 36 CFR part 1194, 28 CFR part 35 or FAR subpart 39.2. Repair suggestions appear in WCAG-EM only as something a commissioner “may request.” Priorities and deadlines appear in enforcement agreements rather than in report formats: the OCR agreement with Santa Clara University requires a Corrective Action Plan setting out “a detailed schedule for addressing problems, taking into account identified priorities.” That obligation runs to the entity that signed the agreement, not to your vendor.