Accessibility Testing

Specify an accessibility audit: scope, sample frame and report

David LoPresti By David LoPresti August 6, 2026

You wrote two paragraphs. Audit our website against WCAG 2.1 Level AA, deliver a report with prioritized recommendations, quote a fixed price. Say four proposals come back: one prices 25 pages, one prices 60 templates, one prices “the full site” with a footnote about crawling, and one asks how many states each form has and quotes three times the lowest bidder. Nothing in them tells you whether they priced the same job.

They have not. An audit is a sample of a product, evaluated against a fixed version of a standard, reported in a fixed shape, and checked by you before you pay. A specification settles three things in that order: the sample frame, which decides what gets looked at and therefore what it costs; the deliverable schema, which decides whether findings are actionable and comparable across bidders; and the acceptance clause, which decides what you are entitled to refuse to pay for.

The substance is already written down by people who do not sell audits: a W3C evaluation methodology, sizing guidance and contract clauses from the General Services Administration, a technical baseline from the Access Board. What follows is quoted from those documents with their disagreements left visible, because two prescribe different sample sizes and neither acknowledges the other. Nothing below needs a particular vendor’s method, ours included. The timing matters: the ADA Title II compliance dates of April 26, 2027 and April 26, 2028 at 28 CFR 35.200(b) fall due inside the next two budget cycles, so if you have never commissioned an audit, you are the one writing the specification anyway.

Name a dated methodology, because “WCAG-EM” changed meaning in July 2026

The Website Accessibility Conformance Evaluation Methodology is the document a scope clause points at, and its 2.0 version was published on 23 July 2026. The W3C Web Accessibility Initiative confirms the date and adds the sentence that matters for drafting: “The 1.0 version of WCAG-EM is still available.”

The two coexist and are not interchangeable: “WCAG-EM 1 was specifically for testing websites and web pages,” while “WCAG-EM 2 also applies to apps and other digital products.” Neither says the newer supersedes the older; 2.0 says only “This document builds on WCAG-EM 1.0.”

Here is the trap. The version-less shortname https://www.w3.org/TR/WCAG-EM/ now serves the 2.0 document, and the 2014 note itself printed that same shortname, then still on http, as its “Latest version”. A clause citing “WCAG-EM” at that URL in June 2026 pointed at one document and now points at another. Name the version and the dated URL.

Comparison of WCAG-EM 1.0 and WCAG-EM 2.0. WCAG-EM 1.0 is the 2014 note and was specifically for testing websites and web pages; WCAG-EM 2.0 was published on 23 July 2026 and also applies to apps and other digital products. Neither document says 2.0 supersedes 1.0; 2.0 says only that it builds on WCAG-EM 1.0, and the 1.0 version is still available. The 2014 note printed the version-less shortname as its own Latest version, and that same shortname now serves the 2.0 document.
Both versions are published and neither retires the other, so a scope clause has to name the version and the dated URL.
View the data as a table
WCAG-EM 1.0WCAG-EM 2.0
PublishedThe 2014 note23 July 2026
What it testsWebsites and web pagesApps and other digital products
Relationship statedStill available; neither document says 2.0 supersedes itSays only that it builds on WCAG-EM 1.0
The version-less shortnamePrinted it as its own Latest versionIs what that shortname serves now

Fix the WCAG version the same way, because “the latest version” is not a specification. Section 508 sits on WCAG 2.0: 36 CFR part 1194, appendix A, provision E205.4, requires that “Electronic content shall conform to Level A and Level AA Success Criteria and Conformance Requirements in WCAG 2.0 (incorporated by reference, see 702.10.1).” That parenthetical is what pins Section 508 to the 2008 Recommendation. Title II sits on WCAG 2.1.

The mismatch is not theoretical. The Access Board’s ICT Testing Baseline for Web still carries a section for Success Criterion 4.1.1 Parsing, and explains why: “Section 508 is not directly affected by WCAG 2.2 as it incorporates by reference WCAG 2.0 Level A and AA, W3C Recommendation, December 11, 2008. SC 4.1.1 Parsing is not deprecated in WCAG 2.0, and the criterion is a Section 508 requirement.” The Baseline then resolves it by adopting the WCAG 2.0 Errata, “This criterion should be considered as always satisfied for any content using HTML or XML,” so its test instruction reads “No testing necessary.” A criterion that lives in one version of the standard and not another got a documented answer here. A floating clause in your contract will not have one.

What a sampled audit cannot give you, at any price

WCAG-EM 2.0 states the ceiling on its own output:

“WCAG 2 conformance claims cannot be made for entire websites based upon the evaluation of a selected sub-set of web pages and functionality alone, as it is always possible that there will be unidentified conformance errors on these websites.”

If your acceptance criterion reads “the auditor shall deliver a WCAG 2.1 Level AA conformance claim for the website,” you have asked the methodology for the one output it says it cannot produce. Two honest alternatives exist. Buy an evaluation of everything, which the methodology prefers where it is achievable: “If feasible, it is recommended to evaluate the entire digital product. The sampling procedure may then be skipped.” Or buy a sampled evaluation and write acceptance against the sample, the method and the findings instead of against a claim. Note also what you are citing: WCAG-EM 2.0 was published “as a Group Note using the Note track” and “does not in any way add to or change the requirements defined by the normative WCAG 2 standard.” It binds a supplier who agreed to follow it, and imports no legal obligation.

The sample frame is the specification

This is the half buyers leave out, and the half that sets the price. WCAG-EM builds it in three moves: define the scope, explore the product, then select.

Define the scope so that every view is in or out. Methodology Requirement 1.1: “Define the target digital product according to Scope of applicability, so that for each view it is unambiguous whether it is within the scope of evaluation or not.” It says how to write it, too: “Using formalizations including regular expressions and listings of web addresses (URIs) is recommended where possible.” A subdomain, a booking widget, a legacy help center and a logged-in area are four separate scope decisions.

Then do not carve pieces out of what you scoped. Under the principle of product enclosure, “Full product enclosure is essential, meaning that we define the scope to include all views, states and functionality of a digital product, without excluding specific parts.” A smaller budget buys a narrower target product, not the same product with the awkward parts removed.

Hand over the exploration outputs or pay for them. Methodology Requirement 3.1: “Select samples that reflect all identified (1) common views, (2) essential functionality, (3) types of samples, (4) technologies relied upon, and (5) other relevant samples.” Two of the five have glossary definitions that settle arguments: common views are “views that are relevant to the entire digital product,” and essential functionality is “functionality that, if removed, fundamentally changes the use or purpose of the product for users.” Write those two lists for your own product and the proposals become comparable. If you cannot, discovery belongs in the statement of work as a priced task.

Complete processes are counted whole. Methodology Requirement 3.3: “Include all samples that are part of a complete process in the selected sample set.” A checkout is not one page. The methodology asks for the default sequence, which “assumes that there are no user input errors and no selection of additional options,” and for the branch sequences “critical for the successful completion of the process.” Error states and alternate paths are named work, and a URL list will not capture them: “In most cases the web address (URL) will not be sufficient to identify the sample in a complete process.”

The random set is 10 percent of the structured set, and it exists to test the frame. Methodology Requirement 3.2 says only “Select a random sample set, and include them for evaluation.” The sizing rule and its worked example sit in the body of Step 3.2:

“The number of samples to randomly select is 10% of the structured sample set selected through the previous steps. For example, if the structured sample set selected for a digital product resulted in 80 samples, then the random sample set size is 8 samples (which are added on top, so in that case, it would leave you with 88 samples in total).”

Its purpose is not extra coverage. It “acts as an indicator to verify that the structured sample set selected through the previous steps is sufficiently representative.” The draw “need not be selected according to strictly scientific criteria,” but it has to span “the entire scope of the digital product” and must not “follow a predictable pattern.” Make the recorded method a deliverable.

Then the comparison, the quality gate built into the methodology. Methodology Requirement 4.3: “Check that each sample in the randomly selected sample set does not show types of content and outcomes that are not represented in the structured sample set.” When that check fails, the methodology sends evaluators back to Step 3 to select additional samples reflecting the newly identified content types and findings. Read that twice before signing a fixed price. It tells the evaluator to enlarge the sample, and says nothing about who pays for the enlargement.

The five parts of a sample frame under WCAG-EM. First, scope defined per view, Methodology Requirement 1.1, so that every view is unambiguously in or out. Second, full product enclosure, with no parts carved out of the scoped product. Third, common views and essential functionality, Methodology Requirement 3.1, together with sample types, technologies relied upon and other relevant samples. Fourth, complete processes counted whole, Methodology Requirement 3.3, covering the default sequence and the critical branch sequences. Fifth, a random sample set at 10 percent of the structured set, Methodology Requirement 3.2, followed by the Methodology Requirement 4.3 comparison back against the structured set.
Each part is a separate drafting decision, and the proposals only become comparable once the buyer has written all five down.
View the data as a list

The sample frame: The half that sets the price

  • Scope per view: MR 1.1: every view in or out
  • Full enclosure: Nothing carved out of scope
  • Common views: MR 3.1: and essential functionality
  • Complete processes: MR 3.3: every page, plus branches
  • Random 10 percent: MR 3.2, then the MR 4.3 check

How many pages should an accessibility audit cover: two published answers

Ask a salesperson and you get a number. Ask the documents and you get two methods that produce different numbers for different reasons.

WCAG-EM does not size the structured set at all. This is the whole of what it says: “The actual size of the sample set needed to evaluate a digital product depends on many factors, including the following:” followed by seven factor headings, “Size of the digital product,” “Age of the digital product,” “Complexity of the digital product,” “Consistency of the product,” “Adherence to development processes,” “Required level of confidence,” and “Availability of prior evaluation findings.” No floor, no ceiling, no ratio to the size of the product. The methodology fixes the ratio of random to structured at 10 percent and leaves the structured number to the evaluator’s judgment, so a proposal that argues that number from the seven factors is doing what the methodology asks. The ratio is not a 2026 development either: WCAG-EM 1.0 carried it in 2014 in almost the same words.

GSA sizes the whole sample statistically, against a population. Its guidance on sample versus comprehensive Section 508 conformance testing states that “Cochran’s formula is used to estimate the number of items that should be tested to achieve a desired level of statistical confidence and precision,” and works it: at a 95 percent confidence level with a plus or minus 5 percent margin of error and p = 0.5, GSA puts “the initial sample size” at “385 (rounded up).” With a known population of 1,000 web pages the number falls to “278 webpages.” One input comes first: “Population size is the total number of assets you could test.” The same page publishes sizes by population, which is the fastest sanity check on a proposal.

Total ICT (population size)Sample size for 90% confidenceSample size for 95% confidenceApproximate margin of error
502630plus or minus 10 percent
1004149plus or minus 10 percent
5007081plus or minus 10 percent
1,0007388plus or minus 10 percent
5,000+90 to 100100 to 150plus or minus 8 to 10 percent

Source: GSA, Sample vs Comprehensive Section 508 Conformance Testing, Table 4, reviewed May 2026. The curve flattens above 5,000 assets, where “the sample size increases slowly.”

Read the last column before using the table as a benchmark. Table 4 is computed at plus or minus 10 percent, not the plus or minus 5 percent of the worked example above, which is why the same population of 1,000 yields 88 here and 278 there. The margin of error, not the standard, is what sets the price. A vendor quoting 88 items and a vendor quoting 278 are both following GSA, and only one of them is selling you the precision you thought you asked for.

Side by side, the disagreement between the two methods is structural rather than cosmetic.

WCAG-EM 2.0, Step 3.2GSA, sample versus comprehensive testing
Status of the documentW3C Group Note, 23 July 2026. Informative. Endorsed by the Accessibility Guidelines Working Group, not by W3C or its MembersUS federal agency guidance page, reviewed May 2026. Not a regulation
What it sizesThe random sample set onlyThe whole sample
The rule10 percent of the structured sample set, added on topCochran’s formula with finite population correction, p = 0.5
Inputs it needsThe structured sample set, whose size it does not prescribePopulation size, confidence level, margin of error
Worked exampleStructured 80, random 8, total 88N = 1,000 at 95 percent and plus or minus 5 percent gives 278; large or unknown population gives 385
What the sample is forVerifying the structured set is representative. A mismatch sends the evaluator back to Step 3Making a statistical inference about the population at a stated confidence and margin of error
What it will not give youA WCAG 2 conformance claim for the whole productA guarantee. GSA describes sampled results as estimates rather than certainties, and says confidence rests on the sample size and how it was selected
Who it bindsNobody by force of law. It binds a supplier who claims to follow WCAG-EMNobody by force of law. It is what a federal buyer is told to do

A third framing comes from GSA as well. Its guidance on including Section 508 in quality assurance surveillance plans has the agency review “a representative sample of webpages (for example, 10% or a risk-based sample) in each sprint or release,” a third 10 percent, of the product rather than of a structured set. None of the three documents cites the others.

Which binds you? If you are a private commercial buyer, none. Pick one, name it, and require the arithmetic in the proposal. Worth importing whichever you choose are GSA’s selection criteria, which prioritize “high traffic and high risk products” and “critical functions and user paths,” ensure content types and components are covered, and close with an instruction that is really a deliverable: “Clearly record how you selected the sample and what confidence level and margin of error apply.”

What counts as one page, and what drops out for software and documents

A page count means nothing until the unit is defined. For a federal buy the Revised 508 Standards define “Web page” at provision E103.4 as “A non-embedded resource obtained from a single Universal Resource Identifier (URI) using HyperText Transfer Protocol (HTTP) plus any other resources that are provided for the rendering, retrieval, and presentation of content.” An application that swaps its whole interface without changing the URI does not divide into that unit, so specify views and states alongside pages.

Two carve-outs in the same appendix reshape a software or document audit. Under E207.2 Exception 3, “Non-Web software shall not be required to conform to Conformance Requirement 3 Complete Processes in WCAG 2.0,” and Exception 2 releases non-web software from 2.4.1 Bypass Blocks, 2.4.5 Multiple Ways, 3.2.3 Consistent Navigation and 3.2.4 Consistent Identification. The exception to E205.4 releases non-web documents from the same four. So the WCAG conformance requirement for complete processes does not reach a desktop application under Section 508. That is not the same as dropping complete processes from the sample. WCAG-EM 2.0 applies to apps as well as websites, and its Methodology Requirement 3.3 still asks for every view in a process. The 508 exception changes what conformance means, not what the evaluator has to look at. If your statement of work covers both, say which rule governs which artifact.

The deliverable schema: outcomes, mappings and machine-readable export

State one thing before writing this clause: no standards body publishes an issue-record schema for an audit report. What exists is a normative vocabulary for test outcomes, a normative field list for mapping a rule to a requirement, an optional recommendation on machine-readable export, and agency guidance on severity. Assemble the schema from those four, each part attributed.

The vocabulary is the strongest piece because it comes from a Recommendation rather than a note. ACT Rules Format 1.1, a W3C Recommendation of 5 February 2026, fixes five outcomes and no others: inapplicable, passed, failed, cantTell, which covers both halves of not knowing, “Whether the rule is applicable, or whether all expectations were met could not be fully determined by the tester,” and untested, “The test subject was not evaluated for the rule.” It also guarantees coverage: “each test subject always has one or more outcomes.”

Require that vocabulary and two kinds of ambiguity leave the report. A criterion nobody looked at reports as untested rather than passing quietly, and one the tester could not resolve reports as cantTell rather than as a pass or a defect. GSA makes the same point about vendor conformance reports: a note of “not evaluated,” it advises, “does not provide any assurance of accessibility.”

The same Recommendation fixes what a mapping must carry: the requirement’s “name, title, identifier or summary,” the name of and a link to “the accessibility requirements document,” the conformance level where one exists, and “whether the requirement is a conformance requirement or a secondary requirement.” Note whose bar that is. Section 4.4 governs the requirements mapping of an ACT Rule, not a finding row in an audit report, so a buyer who wants a finding that says “fails 1.4.3” to name and link its requirements document has to import the field list deliberately. Import it and a finding becomes checkable against a document you both named, instead of a bare criterion number. Rules must also be versioned “with either a date or a number,” and the identifier “must not be changed when the rule is updated,” which is what makes the second audit comparable with the first.

On machine-readable export, be precise about status. WCAG-EM Step 5.5 says “It is recommended to use EARL for providing machine-readable reports.” That step is marked optional, the EARL 1.0 Schema is a W3C Working Group Note from 2 February 2017 rather than a Recommendation, and the JSON-LD examples that would make it operational sit in an ACT appendix labeled “This section is non-normative.” You can require EARL, or CSV, or JSON with a named field list. You cannot say a standard requires it. GSA notes the practical reason in its guidance on the essential elements of an accessibility test report: machine-readable formats “facilitate comparison capabilities and tool integration.”

Severity is yours to define, and the most citable federal definition is functional rather than a set of band names. GSA’s QASP language sets the threshold at “no accessibility defects that prevent users with disabilities from completing required tasks,” treats minor defects as “cosmetic or low-impact issues that do not prevent users with disabilities from being able to perceive, operate, or understand the content,” and settles who decides: “The Government will determine the severity of accessibility defects.” Copy both halves. If the auditor’s scale governs acceptance, the auditor is grading its own work.

Two clauses close the schema. WCAG-EM asks reports to “include at least one example for each conformance requirement and WCAG 2 Success Criterion not met” and to “indicate issues that occur repeatedly.” GSA sets a floor: results for all applicable standards, coverage of “all major features, functions, and product workflows,” and the “testing methodologies and tools used.”

The five fields an audit report issue record has to carry, and where each comes from. Test outcomes: the five ACT Rules Format 1.1 values inapplicable, passed, failed, cantTell and untested, from a W3C Recommendation. Requirements mapping: name, title or identifier, the requirements document and its link, the conformance level, and whether the requirement is a conformance or secondary requirement, from the same Recommendation. Rule identifiers must be versioned with a date or a number and must not change when the rule is updated. Machine-readable export: recommended in WCAG-EM Step 5.5, but that step is optional and EARL is a 2017 Working Group Note. Severity: defined by GSA by effect on task completion, and determined by the government rather than the auditor.
No standards body publishes an issue-record schema, so the buyer assembles one and attributes each field.
View the data as a list

One issue record: Each field attributed to its source

  • Test outcome: One of five ACT values
  • Requirement mapping: Name, document, link, level, type
  • Rule identifier: Versioned, and never changed
  • Export format: Step 5.5 suggests EARL, optional
  • Severity: Blocks a task or not, you decide

Acceptance: the clause, not the review

Name the absence first. WCAG-EM contains no acceptance protocol for the party commissioning the evaluation. It tells the evaluator what to define, select, evaluate and report, and never tells the buyer how to check the evaluator. That half comes from procurement guidance, and the fullest published set is GSA’s, which carries its own caveat: the language is offered “Though not legally required under the Revised 508 Standards.”

Three things go in the clause. The artifacts. GSA’s sample contract language asks that before acceptance the vendor provide a Supplemental Accessibility Conformance Report with test results “based on the required test methods,” the features that aid accessibility, and “Documentation of core functions that cannot be accessed by persons with disabilities.” That last item is the one buyers forget to demand and the one that tells them most. Your own right to test. GSA reserves it before acceptance, “to validate that the ICT solution provided by the contractor conforms to the applicable Revised 508 Standards,” and again in its solicitation guidance, where the agency may “perform testing on some or all of the vendor’s proposed ICT items to validate Section 508 conformance claims made in the ACR.” Both describe the buyer testing the product. No source here describes a buyer re-testing the auditor’s individual findings, and none names a size for the buyer’s verification sample, so both are your own commercial terms rather than cited practice. The payment gate. GSA asks for “a fully working demonstration of the completed ICT Item” that will “expose where such conformance is and is not achieved,” which for an audit means a live walkthrough of findings in the product with the assistive technology named. The QASP language does the rest: “The agency will re-test to confirm that corrective action meets Section 508 requirements,” remediation happens “at no additional cost to the agency,” and acceptance follows the quality level being met.

What the clause buys you is the standing to say no. Which fields make a delivered report refusable on its face, and which ones a vendor can leave out without breaking any methodology, is a separate reading job, worked through here.

What none of these sources settles

Who pays when the sample turns out to be too small. WCAG-EM sends the evaluator back to Step 3 for more samples when the random set surfaces new content types or findings. No source says whether that work sits inside the fixed price or becomes a change order. Settle it in one sentence up front.

The federal buyer’s evaluation criteria are unfinished, by the government’s own admission. On a page reviewed in March 2026, GSA states: “Sample evaluation criteria for accessibility are still being developed.” Writing a factor to score competing audit proposals is not filling in a form that already exists.

Nothing reconciles the sizing methods. Three answers, two authorities, no cross-references.

No published decision turned up. No litigated or administratively decided case on the adequacy of an audit’s sample frame or statement of work appeared in the research for this article, and none of the sources cited here refers to one.

The 2027 and 2028 dates rest on an interim final rule. The Department of Justice document that moved them is published at 91 FR 20902 with the action line “Interim final rule; request for comments,” and its comment period closed on 22 June 2026. No final rule has followed as of publication, so a schedule pinned to April 26, 2027 is pinned to a date the Department invited comment on.

A baseline is a coverage floor, not a procedure. The ICT Testing Baseline for Web says of itself that it “is not intended to be a test process itself” and “does not identify testing tools,” while requiring that “all test processes that claim to align to this baseline must include all baseline tests and provide baseline test results.” Citing it tells you what has to be covered, not how or with what.

What to write this week

Open the draft and do four things in order.

  1. Fix the citations. Standard, version and level; methodology, version and dated URL. Then the target product defined per view, with URI listings or regular expressions where possible, and nothing carved out.
  2. Write the frame. Common views, essential functionality, sample types, technologies relied upon, other relevant samples, and every complete process by name with its default and branch sequences. Supply the lists or price the discovery.
  3. Name the sizing rule and require its arithmetic. Either the seven WCAG-EM factors and the structured number they produce, or a population, confidence level and margin of error. Add the random set at 10 percent, its selection method recorded, and the Step 4.3 comparison reported.
  4. Fix the report and the gate. Five ACT outcomes, the mapping fields, a versioned rule identifier, one example per criterion not met, an export format, a severity scale defined by effect on task completion and determined by you, and an acceptance clause that says what you test yourself, what blocks payment, and who pays for a sample that had to grow.

Federal buys add two paragraphs from outside the audit, each narrower than it first looks. On an indefinite-quantity contract, FAR 39.203(b) requires that “The contract must identify which supplies and services the contractor indicates as compliant and show where full details of compliance can be found,” which is what stops an IDIQ from burying scope across later orders. And 39.203(f), captioned “Alterations of legacy ICT,” closes the grandfathering: “When altering any component or portion of existing ICT, after January 18, 2018, the component or portion must be modified to conform to the current ICT accessibility standards in 36 CFR 1194.1.” That is a conformance obligation on the altered component, not a testing one. The FAR does not supply a retest clause, so if you want the changed portion re-audited, write that yourself.

Where this stops

All of the above concerns test scope and evidence: what gets sampled, how findings are recorded, what you inspect before accepting a report. The questions next door belong to other professions. Whether a clause is enforceable, how liability is allocated, and whether your organization is covered by a given rule at all are for your counsel and, on a federal buy, your contracting officer. Nothing here is legal advice.

If you want the specification written against your product rather than a template, that is what our WCAG audits and testing work covers, and for federal and federally funded buyers the same frame is built against the Revised 508 Standards in our Section 508 compliance work. Related reads: deciding whether you will accept a supplier’s test evidence, for the report you are handed rather than the one you commission; accessibility consulting engagements compared; and kiosk accessibility RFP requirements for hardware buys.