VPAT ACR

We scored 28 published ACRs on documentation quality alone

David LoPresti By David LoPresti July 23, 2026

What this study measures

You have a vendor’s Accessibility Conformance Report open, an award schedule that has not moved, and an instinct that the document is thin. The question is whether that instinct is defensible in writing. This study exists to give you a base rate, so that “this report is unusually weak” becomes a sentence you can put in the file next to a count.

One boundary before any number: this study measures the document, not the product. No product in the sample was tested. Nothing here says whether a vendor’s software conforms to anything. It says what a reviewer can establish by reading the report on the screen, without a license to the product and without a testing budget.

We coded 28 published ACRs against the twelve-flag rubric already published at score a vendor’s ACR, reused without modification. That article scored two reports as a worked example. This is the larger sample on the same instrument.

What has and has not been measured before

The prior art is a 2015 peer-reviewed study by Laura DeLancey at Western Kentucky University, published in Library Hi Tech, which compared 17 VPATs against automated scans of the same products. Across 189 VPAT checkpoints she scanned, 19.6 percent carried inaccurate information. That study measured document against scan. This one reads documents only, which makes the two complements rather than competitors. DeLancey’s figure is not a finding of ours and does not transfer to our sample.

The methodological model is the WebAIM Million: a stated frame, a named tool, a fixed window, and a published statement of what the method cannot see. WebAIM prints the limitation on the project page itself. “Absence of detected errors does not indicate that a page is accessible or conformant.” The equivalent limitation here is the scope sentence above, and it is load-bearing.

One other instrument exists and is worth naming: the University of Central Florida Center for Distributed Learning publishes a VPAT Evaluator that reviews the document rather than the product, and states that it is not a substitute for a comprehensive accessibility audit. It publishes no scoring criteria. That is the gap this pair of articles closes: a rubric anyone can read, and a sample anyone can walk.

Table comparing three instruments on three dimensions. DeLancey 2015 read 17 VPATs against automated scans of the same products, reported that 19.6 percent of 189 scanned checkpoints carried inaccurate information, and was peer reviewed in Library Hi Tech. The UCF VPAT Evaluator reads the document rather than the product, reports that it is not a substitute for a comprehensive accessibility audit, and publishes no scoring criteria. This study reads 28 published ACRs as documents only, reports a base rate for this sample alone because DeLancey's figure does not transfer, and publishes a rubric anyone can read with a sample anyone can walk.
The prior art beside this study. DeLancey’s 19.6 percent is a document-versus-scan figure and does not transfer to this sample. Sources: DeLancey, Library Hi Tech (2015); UCF Center for Distributed Learning VPAT Evaluator.
View the data as a table
DeLancey 2015UCF VPAT EvaluatorThis study
What it reads17 VPATs against automated scans of the same productsThe document rather than the product28 published ACRs, documents only
What it reports19.6 percent of 189 scanned checkpoints carried inaccurate informationThat it is not a substitute for a comprehensive accessibility auditA base rate for this sample only, since DeLancey’s figure does not transfer
Scoring criteriaPeer reviewed, published in Library Hi TechNone publishedA rubric anyone can read, and a sample anyone can walk

Method, published before the results

ElementValue
FrameThe public Library Vendor VPAT Repository published by Houston City College Libraries, a buyer-maintained list of the vendor VPATs and ACRs behind its own subscriptions. Read in full and parsed from the page HTML.
Rows in the frame41
Retrieval and coding windowJuly 27, 2026, single day
IncludedA row that resolved to a readable, VPAT-based ACR: 28
Excluded13, each for a stated reason, listed below
Publishers18, counted by the publisher name printed on the report, not by corporate group
InstrumentThe twelve-flag rubric published at /blog/score-a-vendor-acr, unchanged
CodersOne. A published fire rule was applied to the single judgment-dependent metric. No inter-coder agreement statistic is reported, because a second coder would be needed to produce one.
Date ruleWhere a report gives only a month and year, the first day of that month is used. Age is measured to July 27, 2026.
Reporting ruleCounts printed beside every percentage. The judgment-dependent metric is reported as a floor.

The frame is a buyer’s own collection rather than a vendor’s, and it was not assembled by ADACP, which removes the obvious objection that the sample was picked to produce a result. It also fixes the sector: these are library and scholarly-content products bought by an academic library. The base rates below describe that market. They are not a claim about federal enterprise software, and a reader in another sector should treat them as a method to copy rather than a number to quote.

The exclusions are a finding in their own right, because they describe what a buyer actually finds when it goes back to its own conformance file.

Reason for exclusionRows
Row present in the repository with no link at all1
Link returns 4041
Link returns 200 and serves a login page, so the report is not publicly readable1
Link returns 403 to a normal browser request, so the document could not be read1
Report held in a script-only viewer and not retrievable as a file1
An accessibility statement rather than a VPAT-based ACR4
A Section 508 compliance statement rather than a VPAT-based ACR1
A vendor help-center hub, finder tool or VPAT directory rather than a report3
Total excluded13

Availability was coded on the response body, not on the status code. One row returns HTTP 200 and serves a support-portal login page. A link check alone would have counted that as an available ACR.

Two disclosures that change how the percentages should be read. First, the 28 reports come from 18 publishers: one publisher supplied five reports, one supplied four, three supplied two each, and 13 supplied one each. The sample is 28 documents, not 28 independent authoring practices, and the five-report publisher’s house style moves several counts by itself. Second, one corporate group is already visible in two of the 18 publisher names, covering three of the 28 reports. No systematic parent-company check was run beyond what the reports print, so 18 is a count of names on reports.

Method box for the study. The frame is a buyer's own public VPAT repository of 41 rows, read in full and parsed from the page HTML. It splits three ways: 28 included, the rows that resolved to a readable, VPAT-based ACR; 13 excluded, for a dead link, a login wall, or a document that is not a VPAT-based ACR; and 18 publishers behind the 28, one of which supplied five reports, so a single house style moves several counts. The instrument is the twelve-flag rubric, unchanged, applied by one coder on July 27, 2026, a single day.
The method, published before the results: 41 rows in the frame, 28 readable ACRs from 18 publishers, coded in one day on an unchanged twelve-flag rubric.
View the data as a list

41 rows in a buyer’s public repository: Read in full and parsed from the page HTML

  • Included: 28 ACRs: Rows that resolved to a readable, VPAT-based ACR
  • Excluded: 13 rows: Dead link, login wall, or not a VPAT-based ACR
  • Publishers: 18: One supplied five, so a house style moves counts

Reports are identified below by product category and the report’s own publication month and year. No vendor or product is named. The reason is narrow: this study did not test any product, and pairing a vendor name with a documentation defect invites a reader to hear a conformance verdict that was never made. The frame is public and linked above, so anyone can retrieve the same 41 rows and rebuild the sample; what they cannot do from this article is map a pseudonym onto a name without redoing the coding, which is the intended result.

Terminology, because it decides what counts as a defect

A VPAT is the blank template. ITI states the distinction itself: “Once completed, the VPAT® with documented testing results is referred to as an Accessibility Conformance Report (ACR) that details the accessible features of the tested product or service.” Every document in this sample is an ACR. A defect below is a defect in a report, never in a product.

Result 1: template currency and edition fit

Three of the 28 reports (11 percent) are built on ITI’s current revision. Eighteen are on 2.5. Two label themselves “2.5INT”. One is on 2.4Rev, three on 2.4, and one on 2.3. The dates in the table come from ITI’s own change-tracking file, which lists a date for every 2.x revision.

Template version stated on the reportReportsShare
2.5Rev, ITI’s current revision311%
2.5, October 20231864%
“2.5INT”, not an ITI version string27%
2.4Rev, March 202214%
2.4, February 2020311%
2.3, December 201814%

Read this as a currency finding and nothing more. Section508.gov’s sell-side guidance for vendors names 2.4 as the current version of the template and adds that “Any VPAT® 2.x is acceptable”, which is also a reminder that federal guidance can lag ITI: that page names 2.4 as current while ITI publishes only 2.5Rev files. A rejection on version number is indefensible in a debrief. Note also that ITI’s two dates for the current revision differ. The download listing reads “VPAT 2.5Rev INT (April 2025)” while the change-tracking file inside dates 2.5Rev to January 2025. Attribute whichever you use.

Only “2.5INT” is a genuine labeling error, and a small one: it fuses the revision number with the INT edition, which are two different fields.

Edition is the field that carries weight, because it decides which obligations the report answers at all.

Edition named on the reportReportsWhat it means for a federal buy
International (INT)21Accepted. One of these names its edition “International Criteria Edition”, which is not one of ITI’s four edition names.
Revised Section 5081Accepted
WCAG5The wrong instrument. Section508.gov tells vendors: “If you are selling to the U.S. federal government, then you must use the Revised Section 508 or the INT International Editions of the template”.
No edition named1Cannot be determined from the document

The five WCAG-edition reports are the same five reports that contain no 302.x Functional Performance Criteria table and no 602.x support-documentation table. That is not a coincidence, it is what the edition choice does: none of the five carries a Section 508 chapter table at all. Against a federal solicitation, those five would fire flag 2 on the rubric before anyone reads a conformance row. The buyer holding this frame is a community college rather than a federal agency, so for its own purposes a WCAG-edition report may answer its question. If the same product were offered to a federal agency, the report on file would not, and the missing Chapter 3 rows are the ones that matter most for products whose functions the software chapter does not reach. That mechanism is set out in the Functional Performance Criteria rows an ACR has to answer.

Result 2: how old the evidence is

Seven of the 28 reports (25 percent) carry a report date more than 24 months before July 27, 2026. The oldest is dated April 2019, which makes it 87 months old.

Age of the report at July 27, 2026ReportsShare
Under 12 months829%
12 to 24 months1346%
24 to 36 months414%
Over 36 months311%

The median report in the sample is 18.7 months old. Three of the seven sit between 24 and 26 months, so the threshold does real work: on a 36-month rule the count would be three rather than seven. Publish your threshold in the solicitation and the argument disappears.

One report in the 28 has no Report Date element at all. It supplies a “Completion Date” and a “Revision Date” under a heading structure that is not the template’s. The template sets the floor plainly. Its Essential Requirements list the Report Date element as the “Date of report publication”, with the instruction “At a minimum, provide the month and year of the report publication.” That report is also the one naming a non-ITI edition, and it is the library discovery website, June 2024.

Age is not the same as the rubric’s flag 3, which asks whether the report date and product version match the release being offered. That cannot be coded from a document alone, because it needs the solicitation and the offered build. Age is the part a reviewer can code off the cover page, and it is the part that tells you whether to ask.

Result 3: the required test-method field is populated and still does not answer the question

“Evaluation Methods Used” is one of the template’s minimum content elements, and the instruction attached to it in the Essential Requirements is to “Include a description of evaluation methods used to complete the VPAT for the product under test.” On presence, the sample is close to clean: 27 of 28 carry the element under the template’s own heading, and the 28th carries comparable content under the heading “Evaluation Methods”.

Depth is where the sample separates. Measured strictly inside that element, and counting only the words inside it:

Words inside the Evaluation Methods Used elementReports
No element under the template’s heading1
1 to 14 words10
15 to 100 words10
More than 100 words7

The median is 37 words. The shortest populated element in the sample is three words long: an automated tool name and its version number, with no method, no scope and no assistive technology. The longest runs 387 words and names the standard and levels, two operating systems with their browsers, the tools and two screen readers. That spread is the finding. The same required field, answered by two vendors in the same market, can be a test plan or a product name.

Table comparing the shortest and the longest Evaluation Methods Used elements in the 28-report sample. The shortest runs 3 words, names an automated tool name and its version number, names no assistive technology, and gives the reviewer a product name. The longest runs 387 words, names the standard and levels together with two operating systems and their browsers and the tools, names two screen readers, and gives the reviewer a test plan. The sample median is 37 words.
Both reports populate the same required element. The median across the 28 is 37 words. Source: the Evaluation Methods Used elements of the 28 ACRs coded for this study.
View the data as a table
Shortest populated elementLongest element
Words inside the element3387
What it namesAn automated tool name and its version numberThe standard and levels, two operating systems with their browsers, the tools
Assistive technologyNone namedTwo screen readers
What the reviewer getsA product nameA test plan

Two of the 28 populate the required element with a statement of tester familiarity rather than a method. One reads, in full: “Testing based on general product knowledge.” The other reads “Evaluation is based on general product knowledge and testing of website.” Both are answering the first item on ITI’s list of things an author may enter under the element, which reads “Indicate whether testing is performed by testers with knowledge of general product functionality”, followed by ITI’s own instructional note that this means the tester knows the common uses and flows of the product in addition to accessibility. It is a disclosure about who tested, not a description of what was done. ITI evidently found the phrase confusing too: the change-tracking file records that revision 2.4Rev exists in part to “clearly describe what is meant by ‘Testing is based on general product knowledge’”.

That leaves rubric flag 7 with two prongs to score, absent or non-responsive, and they do not give the same answer.

On the absent prong, flag 7 fires on none of the 28. On the non-responsive prong it is arguable on those two, because the element’s own instruction heading in Best Practices for Authors is “Describe the testing performed”, and neither sentence describes any testing. This study scores both prongs clean, for a stated reason: the published rubric’s gloss on flag 7 is to score the field and not its richness, and a vendor who wrote one of those sentences will reply that it entered the first item ITI lists. That reply survives a debrief. A reviewer who reads “non-responsive” more strictly should score those two reports at zero on flag 7 and let the published evidence cap do the rest, since flag 5, 6 or 7 at zero holds the total at 64 whatever the other eleven flags produce. Either way the useful move is the same, and it is not an argument about the template.

One publisher’s five reports each populate the element with the same ten or eleven word list: product knowledge, a monitoring platform, a screen reader, a browser-based checker and browser developer tools. Five of the 28 counts come from one house style, which is exactly why the publisher concentration is disclosed above.

If you want the version of this question you can actually put to an offeror, the wording is in will you accept our test evidence.

Result 4: assistive technology named, versions absent

Coded fieldReportsShare
Names at least one assistive technology anywhere in the report2175%
Names at least one specific tool or assistive technology inside the Evaluation Methods Used element2382%
Names no specific tool or assistive technology inside that element518%
Names an assistive technology together with its version311%

Three of 28 is the single most reviewable gap in the sample, and it is not a template violation. ITI marks it optional, in the Best Practices for Authors section of the template file itself, in these words: “Describe testing conducted with assistive technologies (Optional: Include the assistive technologies that were used in testing.)” A vendor that omits assistive technology versions has broken nothing. The demand has to come from your solicitation, and citing the template for it is how a reviewer loses an argument with a capture manager.

What the three reports that do it look like: one names three annual releases of a commercial screen reader plus a version of a free one; one names a screen reader build and the browser build it was driven in; one names a screen reader release from 2021 and another from 2022, which is itself informative, because it dates the testing rather than the report. A fourth report gives a tool version and no assistive technology version, which is the three-word element described above.

Screen reader and browser versions are the difference between a claim you can reproduce and a claim you can only believe. Which combinations are worth naming in a solicitation is covered in assistive technology test targets.

Result 5: no Supplemental Accessibility Report anywhere in the sample

Zero of the 28 reports supply a Supplemental Accessibility Report. The word “Supplemental” appears in five of the 28, and every one of those five was read in context: in each case it is ordinary prose about supplemental content, never a report heading.

This one points back at the buyer. The SAR is where the buy-side guidance puts evaluation methods, features that help achieve accessibility, core functions that cannot be used by persons with disabilities, and configuration and installation guidance. Section508.gov’s buy-side guidance is also where the pressure point sits, because it states that “To be considered for award, the ACR must be complete, and submitted according to the instructions.” If your solicitation never asked for a SAR, its absence across 28 reports is a gap in solicitations rather than a gap in vendor practice, and the fix is a clause rather than a rejection. The clause library is in Section 508 contract clauses and the QASP.

The rubric’s flag 10 carries a scoring rule worth restating here: score the content, not the heading. Partial credit depends on how many of the SAR’s four elements survive inside the ACR itself, and this study coded only one of them, the evaluation-methods element, which is present in 27 of 28. So read the zero as “no report supplied the artifact”, not as “flag 10 fires at full weight 28 times”.

Result 6: Supports rows that do not survive their own remarks

This is the one judgment-dependent metric, so the rule comes before the count.

ITI’s definition, which the rubric turns on: “Supports: The functionality of the product has at least one method that meets the criterion without known defects or meets with equivalent facilitation.” The published fire rule, quoted from the rubric: fire only “when no conforming method survives the remark: a defect in the only method available, or a fix scheduled for a future release.” A defect in one of several methods does not fire it. A documented workaround alongside a path that still conforms does not fire it. An exception that the success criterion itself allows does not fire it.

How the metric was produced, in two passes. First, a keyword detector surfaced every Supports row whose remark contained one of a fixed list of defect markers: exception, known issue, known defect, known limitation, will be fixed, will be addressed, will be remediated, will be resolved, will be corrected, future release, next release, upcoming release, roadmap, scheduled, workaround, filed as bug, internal ticket, planned. Second, every detector hit was read in full against the fire rule. The detector over-fired heavily, and no row below was confirmed on a detector hit alone.

Confirmed in 5 of the 28 reports, across 7 rows. Report this as a floor, not an estimate: the detector is keyword driven, so a Supports row describing a defect without any of those markers would have been missed, and one coder applied the rule rather than two.

Do and don't card for scoring a Supports row against its own remark. Do fire the flag only when no conforming method survives the remark, meaning a defect in the only method available or a fix scheduled for a future release; publish the fire rule before the count, since this is the one judgment-dependent metric; read every detector hit in full, because a keyword detector over-fires heavily; and report the confirmed count as a floor rather than an estimate, 5 of 28 reports across 7 rows. Do not fire it on a defect in one of several methods or on a workaround alongside a path that still conforms; do not fire it on an exception the success criterion itself allows, such as confirming passwords at 3.3.7 Redundant Entry; do not confirm a row on a keyword alone, since the phrase no known defects matches the detector on the word defect inside a negation; and do not rescore a report's own declared vocabulary, where 31 Supports with Exceptions rows carry ITI's Partially Supports definition.
The fire rule as published, with what it kept out. Confirmed in 5 of the 28 reports, across 7 rows, and reported as a floor.
View the data as a table
DoDon’t
Fire it only when no conforming method survives the remark: a defect in the only method available, or a fix scheduled for a future release.Fire it on a defect in one of several methods, or on a workaround that sits alongside a path which still conforms.
Publish the fire rule before the count, since this is the one judgment-dependent metric.Fire it on an exception the success criterion itself allows, such as confirming passwords at 3.3.7 Redundant Entry.
Read every detector hit in full against the rule, because a keyword detector over-fires heavily.Confirm a row on a keyword alone: “no known defects” matches the detector on the word defect inside a negation.
Report the confirmed count as a floor rather than an estimate: 5 of 28 reports, across 7 rows.Rescore a report’s own declared vocabulary: its 31 “Supports with Exceptions” rows carry ITI’s Partially Supports definition.
Report (category, report month)CriterionLevel as statedThe remark that contradicts it
Online reference and eBook platform, December 20232.4.6 Headings and Labels (AA)SupportsUnder a heading “Current Exceptions:” a single bullet, “Some labels may be incorrectly applied or missing.”
Online reference and eBook platform, December 20233.1.1 Language of Page (A)SupportsThe remark opens “Supports with the current possible exception:” and then reports that with a third-party translation plugin in use “the lang attribute does not switch to the new translated language”
Ethnographic research database, February 20262.1.1 Keyboard (A)Supports”An exception is noted for the Filters. This is due to a 3rd party library. See conformance and remediation documentation for details on a fix.”
Library service application B, September 20242.4.6 Headings and Labels (AA)Web: Supports, Authoring Tool: Partially Supports”Web: There are some small exceptions that were filed as bugs that we will need to address”, followed by two numbered items naming modal titles set as H4 on pages with no H3, and results information nested in an H2
Library service application B, September 20242.4.1 Bypass Blocks (A)Web, Electronic Docs and Authoring Tool: Supports”Authoring Tool: A small exception is that there is no way right now to keyboard bypass the user list names on the Appointments -> Calendar interface to get to the main calendar content.”
Library service application C, October 20241.4.11 Non-text Contrast (AA)Web: Supports, Authoring Tool: Partially Supports”Web: The Cancel button background on the blog comment section is similar to the blog background color. Users can still find the Cancel button by its context.”
Library service application D, December 20243.2.3 Consistent Navigation (AA)Web, Electronic Docs and Authoring Tool: Supports”Authoring Tool: There is an exception with the widget preview screens. The widget preview screens do not have the common top nav bar.”

The 1.4.11 row is the weakest of the seven and the one a vendor would contest first, so it is printed in full rather than trimmed. The vendor’s second sentence, that users can still find the Cancel button by its context, is a statement about impact, not a second method. Success criterion 1.4.11 is a contrast threshold, and locating a control by context does not meet a contrast threshold or supply equivalent facilitation for it; the remark records a known defect in the only Cancel button on that screen. That is the reading applied here. A reviewer who accepts the vendor’s sentence as sufficient should record 6 rows in 4 reports instead, and either way the action is a question to the vendor rather than a rejection. Note also that the same cell ends with “An internal ticket has been filed for these issues”, which covers the Authoring Tool defect in the same cell, so it is not quoted above as though it attached to the Web row.

What the rule kept out matters as much as what it let in, because these are the rejections that stop the metric from being an opinion.

  • Remarks reading “no known defects” were matched by the detector on the word “defect” inside a negation. Rejected.
  • Rows at 3.3.7 Redundant Entry citing “the allowed exception of confirming passwords during account creation” invoke the success criterion’s own exception. Rejected.
  • Rows at 1.4.5 Images of Text citing graphs and equations that must be preserved for meaning, and rows at 1.1.1 citing purely decorative images, invoke exceptions the criteria themselves carry. Rejected.
  • One report uses “Supports with Exceptions” where the template has “Partially Supports”, declares the substitution in its own Terms section, and gives it ITI’s Partially Supports definition verbatim: “Some functionality of the product does not meet the criterion.” The string appears 32 times in the file, one of which is that Terms definition, so 31 are row-level uses, and “Partially Supports” appears nowhere. Counting those 31 rows as contradicted Supports rows would have taken the headline from 7 rows to 38. Under the report’s own declared definitions the rows are internally consistent, so all 31 were rejected here. This is the sample’s biggest false-positive trap.
  • Two further detector hits in one report were read and not confirmed, and one row asserting that an exception is “valid” without naming a basis was left unconfirmed. It is a question for the vendor, not a finding.

The 31-row case is worth one more sentence, because it is not a vendor inventing vocabulary. “Supports with Exceptions” is ITI’s own retired level. ITI’s VPAT page records that “in a previous update of the VPAT, ‘partially supports’ replaced ‘supports with exceptions.’ This change was made at the request of representatives of the U.S. Access Board”, and the change-tracking file puts that swap in version 2.2, June 2018, for the 508 and INT editions. The report claims VPAT 2.4 on its cover, names no edition, and runs pre-2.2 vocabulary throughout: it also prints “2017 Section 508” 43 times, which is the label 2.2 replaced with “Revised Section 508”. The correct finding is not a coinage. It is a report whose Terms section is three revisions behind the version it claims, which is a template deviation and is recorded as one below.

One report is recorded separately, because the defect is remark specificity rather than a level contradiction: under three Supports levels at 1.3.1 Info and Relationships it announces that a monitoring tool “has found some minor exceptions”, follows the colon with no list at all, and leaves a second scope label in the same cell with no text after it. That is the rubric’s flag 6 territory, and flag 6 is not reported across this sample. Row-level extraction from multi-column PDF tables was not reliable enough to publish: four reports left between 9 and 55 rows with no parseable conformance level, which is almost certainly an extraction artifact rather than a missing level. Publishing a remark-specificity percentage off that would have been an opinion dressed as data.

Result 7: defects visible without reading a single conformance row

Six of the 28 reports carry at least one of the defects below. Five of the six are template deviations, which is rubric flag 12, and the sixth is the “Not Evaluated” case, which is flag 4. The list of defect types was fixed by reading rather than by an exhaustive template audit, so treat both counts as floors.

DeviationReportsWhat a reviewer sees
Blank-template front matter still attached3ITI’s instructions to vendors sit ahead of the report, so the buyer receives the manual and the report in one file. Two of the three also show unresolved Microsoft Word field codes where a table of contents should be. A reader searching for “Report Date” hits the instruction text first.
Conformance level renamed1”Supports with Exceptions”, the level ITI retired in 2018, on 31 rows, declared in the report’s own Terms.
No Report Date element, and a non-ITI edition name1”Completion Date” and “Revision Date” in place of the template’s Report Date.
Product Description describes a different product1The description in one report is the same sentence used in another report from the same publisher, and it describes the other product.
”Not Evaluated” on Level A or AA rows1Six rows, all of them WCAG 2.2-only criteria.

The front matter row and its mirror image are worth reading together, because one of them is not a defect at all. ITI’s instruction to vendors is unambiguous: “When publishing your Accessibility Conformance Report, be sure to remove the entire first 10 pages of this document, including the table of contents, introductory information and instructions.” Section508.gov’s sell-side guidance says the same thing and names the heading to cut to. So three of these 28 shipped pages their vendor was told to delete. A fourth report deleted them correctly and looks worse for it: its page-number field still counts from the untrimmed template, so a complete seven-page ACR opens at “Page 10 of 16” and repeats that pagination to the end. It carries the full heading block, contact information with an email and a phone number, Evaluation Methods Used, Terms, and both the Level A and Level AA tables. Nothing is missing. It is recorded here as a reading trap rather than a deviation, because a reviewer who rejects a report for starting at page 10 will be rejecting the vendor who followed the instructions.

Two of the table rows need their reading stated, or the count gets used to say something it does not support.

The copy-forward Product Description is the cleanest evidence in the sample that a finished document went out unread. It is a documentation defect and nothing more. It says nothing about either product.

The “Not Evaluated” case is a template-form defect, not a vendor declining to test what it claimed. The six rows are 3.2.6 Consistent Help and 3.3.7 Redundant Entry at Level A, and 2.4.11 Focus Not Obscured (Minimum), 2.5.7 Dragging Movements, 2.5.8 Target Size (Minimum) and 3.3.8 Accessible Authentication (Minimum) at Level AA, all of them WCAG 2.2-only criteria, and the same report’s Applicable Standards table declares WCAG 2.2 Level A and AA out of scope. The template restricts the level plainly: “Not Evaluated: The product has not been evaluated against the criterion. This can only be used in WCAG Level AAA criteria.” Either “Not Applicable” or an omitted 2.2 table would have said the same thing in a template-consistent way. Which WCAG version a given US rule actually requires is a separate question, set out in WCAG version requirements by rule.

Flag 12 has a consequence that reviewers underuse. The template’s Essential Requirements for Authors state that “Users of the VPAT agree not to deviate from the Essential Requirements for Authors”, and that using the template and the name requires the registered service mark. A report that deviates materially is not just untidy, it is a document whose author has stepped outside the terms under which the template may be called a VPAT at all. That is a reason to ask for a reissue, and it is a small ask.

Contact information, coded on the template’s own bar that “Listing an email is sufficient”: one report of 28 leaves the Contact Information element empty, with no email address anywhere in its 55 pages. Two more give no reachable route, both reading that the reader should contact an institutional sales representative. That is flag 11 firing once and arguable twice.

Which of the twelve flags a document study can score

The rubric’s twelve flags and weights are reproduced here unchanged, with what this study could and could not code against each.

#FlagWeightCoded hereResult across the 28
1Applicable Standards/Guidelines indication absent, or inconsistent with the edition used8PartlyNot coded across the sample. Two instances found by reading: one report whose 2.2 rows read “Not Evaluated” while its own scope table declares 2.2 out of scope, one report naming a non-ITI edition.
2Federal 508 obligations not reported (a WCAG-only claim against a 508 buy)10YesFires on 5 of 28, the five WCAG-edition reports with no 302.x and no 602.x table.
3Report date or product version does not match the release being offered8NoNeeds the solicitation and the offered build. Age coded instead: 7 of 28 over 24 months.
4”Not Evaluated” appearing on a Level A or Level AA success criterion8YesFires on 1 of 28, 6 rows, all WCAG 2.2-only criteria.
5A “Supports” row whose own remark leaves no conforming method standing12Yes, under the published fire ruleConfirmed in 5 of 28, 7 rows. A floor.
6Thin or missing remarks on “Partially Supports” and “Does Not Support” rows12NoRow-level extraction was not reliable enough to publish. Not reported.
7”Evaluation Methods Used” absent or non-responsive12YesFires on 0 of 28 on the absent prong. Arguable on 2 of 28 on the non-responsive prong, scored clean here for the reason given in Result 3.
8Assistive technologies and testing tools not named4Yes5 of 28 name none inside the element. 3 of 28 name an assistive technology with a version.
9Chapter 3 Functional Performance Criteria not answered where Chapters 4 and 5 leave a function uncovered10Conditional, not codedFiring it requires establishing which product functions the hardware and software chapters do not reach. Table presence coded instead: 23 of 28 carry a 302.x table.
10No Supplemental Accessibility Report8Yes0 of 28 supply one.
11No contact information for follow-up questions4YesEmpty element in 1 of 28. No reachable route in 2 more.
12Material deviation from the template’s essential requirements4Yes, for a fixed list of deviation typesAt least one recorded deviation in 5 of 28.

No total scores and no score distribution are published, and the reason is a rule rather than caution. Flag 3 needs the offered release. Flag 6 needs reliable row-level extraction this sample did not have. Flag 9 is conditional on a product-function finding no document study makes. The rubric also caps any total at 64 when flag 5, 6 or 7 scores zero, so a total assembled from nine flags out of twelve would not be the instrument’s output. It would be a different number wearing the instrument’s name.

That is the honest answer to the question the sample was built to ask. Of twelve flags, seven can be coded from the document on your screen and five cannot. Seven is enough to tell you whether to accept, return, or spend money on testing.

What the document does not tell you

It does not tell you whether the product conforms. A report can be current, on the right edition, specific about method, name assistive technology versions, and still describe a product with defects nobody found. The reverse holds too: a thin report can describe a well-built product whose vendor writes badly. Documentation quality is a proxy for the diligence behind a claim, not a measurement of the claim.

It also does not tell you what a script or widget bolted onto the product achieves, since that claim lives in the product and not in the document.

There is no central federal source to check a vendor’s report against. GSA’s own ACR Library states that reports “are provided only for the tools and training made available by GSA through Section508.gov”. Counted on July 27, 2026 it holds nine rows, three tools and six online training courses, three of them showing a Report Date of “Pending” with no downloadable report. It contains no third-party vendor ACRs. You evaluate what the offeror hands you.

And the verification gap on the buyer’s side is documented, though it has to be quoted precisely. GSA’s FY 2025 governmentwide Section 508 assessment records agency self-report on how often ICT deliverables from a contract are verified for Section 508 conformance: never 7 percent, rarely 20 percent, sometimes 23 percent, often 13 percent, almost always 30 percent, and 7 percent recording that they do not perform the activity. That is a statistic about agency acquisition practice, self-reported, and the assessment states its own limitation: “Agencies, parent agencies, and components self-reported the data in this report. No independent validation or external data was utilized.” It is not a measure of ACR accuracy and it does not mean 30 percent of ACRs go unverified.

What to put in the next solicitation

The three fields this study found thinnest are optional in the template, which means they arrive only if you ask for them in writing. All three cost a vendor a paragraph.

  1. Assistive technologies and tools, with versions. Optional in ITI’s Best Practices, absent with versions in 25 of the 28 reports here. Ask for the assistive technology, its version, the browser and its version, and the operating system. Ask through the Supplemental Accessibility Report, not as a template requirement.
  2. A Supplemental Accessibility Report for each standard commercial item. Zero of 28 supplied one. This is the artifact that carries evaluation methods, features that help achieve accessibility, core functions that cannot be used by persons with disabilities, and configuration and installation guidance, and it is the correct home for the two asks above and below.
  3. A test-method description that names scope. The required element is populated in 27 of 28 and reaches 14 words or fewer in 10 of them. Specify what “described” means: which platforms, how many screens or documents, which standard and level, manual as well as automated, and who performed it.
Ranked list of three asks for the next solicitation. First, assistive technologies and tools with versions: optional in ITI's Best Practices and absent with versions in 25 of the 28 reports, so ask for the version, the browser and the operating system. Second, a Supplemental Accessibility Report: zero of 28 supplied one, and it is the correct home for evaluation methods, unusable core functions and configuration guidance. Third, a test-method description that names scope: the element is populated in 27 of 28 and runs to 14 words or fewer in 10 of them, so specify platforms, screens, standard and level, and manual as well as automated testing.
The three fields this study found thinnest. All three are optional in the template, so they arrive only if you ask for them in writing.
View the data as a list
  1. Assistive technologies and tools, with versions: Optional in ITI’s Best Practices, absent with versions in 25 of the 28. Ask for the version, the browser and the operating system.
  2. A Supplemental Accessibility Report: Zero of 28 supplied one. It is the correct home for evaluation methods, unusable core functions and configuration guidance.
  3. A test-method description that names scope: Populated in 27 of 28, and 14 words or fewer in 10 of them. Specify platforms, screens, standard and level, manual as well as automated.

Then state your own thresholds. Publish the maximum report age you will accept, publish that the edition must be Revised 508 or INT for a federal buy, and publish that completeness is measured against your ACR instructions. Section508.gov’s buy-side guidance already gives you the award condition; your instructions are what make it operable.

The sentence that goes in the file is not “this ACR looked weak”. It is closer to this: this firm applied a published twelve-flag instrument consistently across offerors; against a 28-report sample of published ACRs scored on the same instrument, this report is past the 24-month age threshold that 21 of the 28 met, states its evaluation method in fewer words than the 37-word sample median, and names no assistive technology version, a field 3 of the 28 populate; we are therefore exercising the reserved right to test before award. That is a defensible record whichever way the decision goes. Where several offerors land in the same place on the same field, that is a market finding rather than a vendor finding, and it belongs in the file as one.

Your next step

Take the ACR on top of your pile and code five fields against the base rates above. Template edition, as named on the report. Report date, and its age today. The Evaluation Methods Used element, word for word, with a word count. Whether any assistive technology carries a version. Whether a Supplemental Accessibility Report exists. If the report is over 24 months old, states its method in 14 words or fewer, and names no assistive technology version, it sits with the 7 of 28, the 10 of 28 and the 25 of 28 respectively. Five fields, three counts, and a finding you can write down instead of an impression you have to defend.

If the gap turns out to be on your side of the table, ADACP’s Section 508 procurement support covers ACR review and the solicitation language that produces reviewable reports in the first place, including the Supplemental Accessibility Report clause that every finding above depends on. Vendors reading this from the other side: vendor guidance on ACRs and VPAT testing apply the same twelve flags before a report is published, which is the cheaper end of this exercise.