We scored 28 published ACRs on documentation quality alone
What this study measures
You have a vendor’s Accessibility Conformance Report open, an award schedule that has not moved, and an instinct that the document is thin. The question is whether that instinct is defensible in writing. This study exists to give you a base rate, so that “this report is unusually weak” becomes a sentence you can put in the file next to a count.
One boundary before any number: this study measures the document, not the product. No product in the sample was tested. Nothing here says whether a vendor’s software conforms to anything. It says what a reviewer can establish by reading the report on the screen, without a license to the product and without a testing budget.
We coded 28 published ACRs against the twelve-flag rubric already published at score a vendor’s ACR, reused without modification. That article scored two reports as a worked example. This is the larger sample on the same instrument.
What has and has not been measured before
The prior art is a 2015 peer-reviewed study by Laura DeLancey at Western Kentucky University, published in Library Hi Tech, which compared 17 VPATs against automated scans of the same products. Across 189 VPAT checkpoints she scanned, 19.6 percent carried inaccurate information. That study measured document against scan. This one reads documents only, which makes the two complements rather than competitors. DeLancey’s figure is not a finding of ours and does not transfer to our sample.
The methodological model is the WebAIM Million: a stated frame, a named tool, a fixed window, and a published statement of what the method cannot see. WebAIM prints the limitation on the project page itself. “Absence of detected errors does not indicate that a page is accessible or conformant.” The equivalent limitation here is the scope sentence above, and it is load-bearing.
One other instrument exists and is worth naming: the University of Central Florida Center for Distributed Learning publishes a VPAT Evaluator that reviews the document rather than the product, and states that it is not a substitute for a comprehensive accessibility audit. It publishes no scoring criteria. That is the gap this pair of articles closes: a rubric anyone can read, and a sample anyone can walk.

View the data as a table
| DeLancey 2015 | UCF VPAT Evaluator | This study | |
|---|---|---|---|
| What it reads | 17 VPATs against automated scans of the same products | The document rather than the product | 28 published ACRs, documents only |
| What it reports | 19.6 percent of 189 scanned checkpoints carried inaccurate information | That it is not a substitute for a comprehensive accessibility audit | A base rate for this sample only, since DeLancey’s figure does not transfer |
| Scoring criteria | Peer reviewed, published in Library Hi Tech | None published | A rubric anyone can read, and a sample anyone can walk |
Method, published before the results
| Element | Value |
|---|---|
| Frame | The public Library Vendor VPAT Repository published by Houston City College Libraries, a buyer-maintained list of the vendor VPATs and ACRs behind its own subscriptions. Read in full and parsed from the page HTML. |
| Rows in the frame | 41 |
| Retrieval and coding window | July 27, 2026, single day |
| Included | A row that resolved to a readable, VPAT-based ACR: 28 |
| Excluded | 13, each for a stated reason, listed below |
| Publishers | 18, counted by the publisher name printed on the report, not by corporate group |
| Instrument | The twelve-flag rubric published at /blog/score-a-vendor-acr, unchanged |
| Coders | One. A published fire rule was applied to the single judgment-dependent metric. No inter-coder agreement statistic is reported, because a second coder would be needed to produce one. |
| Date rule | Where a report gives only a month and year, the first day of that month is used. Age is measured to July 27, 2026. |
| Reporting rule | Counts printed beside every percentage. The judgment-dependent metric is reported as a floor. |
The frame is a buyer’s own collection rather than a vendor’s, and it was not assembled by ADACP, which removes the obvious objection that the sample was picked to produce a result. It also fixes the sector: these are library and scholarly-content products bought by an academic library. The base rates below describe that market. They are not a claim about federal enterprise software, and a reader in another sector should treat them as a method to copy rather than a number to quote.
The exclusions are a finding in their own right, because they describe what a buyer actually finds when it goes back to its own conformance file.
| Reason for exclusion | Rows |
|---|---|
| Row present in the repository with no link at all | 1 |
| Link returns 404 | 1 |
| Link returns 200 and serves a login page, so the report is not publicly readable | 1 |
| Link returns 403 to a normal browser request, so the document could not be read | 1 |
| Report held in a script-only viewer and not retrievable as a file | 1 |
| An accessibility statement rather than a VPAT-based ACR | 4 |
| A Section 508 compliance statement rather than a VPAT-based ACR | 1 |
| A vendor help-center hub, finder tool or VPAT directory rather than a report | 3 |
| Total excluded | 13 |
Availability was coded on the response body, not on the status code. One row returns HTTP 200 and serves a support-portal login page. A link check alone would have counted that as an available ACR.
Two disclosures that change how the percentages should be read. First, the 28 reports come from 18 publishers: one publisher supplied five reports, one supplied four, three supplied two each, and 13 supplied one each. The sample is 28 documents, not 28 independent authoring practices, and the five-report publisher’s house style moves several counts by itself. Second, one corporate group is already visible in two of the 18 publisher names, covering three of the 28 reports. No systematic parent-company check was run beyond what the reports print, so 18 is a count of names on reports.

View the data as a list
41 rows in a buyer’s public repository: Read in full and parsed from the page HTML
- Included: 28 ACRs: Rows that resolved to a readable, VPAT-based ACR
- Excluded: 13 rows: Dead link, login wall, or not a VPAT-based ACR
- Publishers: 18: One supplied five, so a house style moves counts
Reports are identified below by product category and the report’s own publication month and year. No vendor or product is named. The reason is narrow: this study did not test any product, and pairing a vendor name with a documentation defect invites a reader to hear a conformance verdict that was never made. The frame is public and linked above, so anyone can retrieve the same 41 rows and rebuild the sample; what they cannot do from this article is map a pseudonym onto a name without redoing the coding, which is the intended result.
Terminology, because it decides what counts as a defect
A VPAT is the blank template. ITI states the distinction itself: “Once completed, the VPAT® with documented testing results is referred to as an Accessibility Conformance Report (ACR) that details the accessible features of the tested product or service.” Every document in this sample is an ACR. A defect below is a defect in a report, never in a product.
Result 1: template currency and edition fit
Three of the 28 reports (11 percent) are built on ITI’s current revision. Eighteen are on 2.5. Two label themselves “2.5INT”. One is on 2.4Rev, three on 2.4, and one on 2.3. The dates in the table come from ITI’s own change-tracking file, which lists a date for every 2.x revision.
| Template version stated on the report | Reports | Share |
|---|---|---|
| 2.5Rev, ITI’s current revision | 3 | 11% |
| 2.5, October 2023 | 18 | 64% |
| “2.5INT”, not an ITI version string | 2 | 7% |
| 2.4Rev, March 2022 | 1 | 4% |
| 2.4, February 2020 | 3 | 11% |
| 2.3, December 2018 | 1 | 4% |
Read this as a currency finding and nothing more. Section508.gov’s sell-side guidance for vendors names 2.4 as the current version of the template and adds that “Any VPAT® 2.x is acceptable”, which is also a reminder that federal guidance can lag ITI: that page names 2.4 as current while ITI publishes only 2.5Rev files. A rejection on version number is indefensible in a debrief. Note also that ITI’s two dates for the current revision differ. The download listing reads “VPAT 2.5Rev INT (April 2025)” while the change-tracking file inside dates 2.5Rev to January 2025. Attribute whichever you use.
Only “2.5INT” is a genuine labeling error, and a small one: it fuses the revision number with the INT edition, which are two different fields.
Edition is the field that carries weight, because it decides which obligations the report answers at all.
| Edition named on the report | Reports | What it means for a federal buy |
|---|---|---|
| International (INT) | 21 | Accepted. One of these names its edition “International Criteria Edition”, which is not one of ITI’s four edition names. |
| Revised Section 508 | 1 | Accepted |
| WCAG | 5 | The wrong instrument. Section508.gov tells vendors: “If you are selling to the U.S. federal government, then you must use the Revised Section 508 or the INT International Editions of the template”. |
| No edition named | 1 | Cannot be determined from the document |
The five WCAG-edition reports are the same five reports that contain no 302.x Functional Performance Criteria table and no 602.x support-documentation table. That is not a coincidence, it is what the edition choice does: none of the five carries a Section 508 chapter table at all. Against a federal solicitation, those five would fire flag 2 on the rubric before anyone reads a conformance row. The buyer holding this frame is a community college rather than a federal agency, so for its own purposes a WCAG-edition report may answer its question. If the same product were offered to a federal agency, the report on file would not, and the missing Chapter 3 rows are the ones that matter most for products whose functions the software chapter does not reach. That mechanism is set out in the Functional Performance Criteria rows an ACR has to answer.
Result 2: how old the evidence is
Seven of the 28 reports (25 percent) carry a report date more than 24 months before July 27, 2026. The oldest is dated April 2019, which makes it 87 months old.
| Age of the report at July 27, 2026 | Reports | Share |
|---|---|---|
| Under 12 months | 8 | 29% |
| 12 to 24 months | 13 | 46% |
| 24 to 36 months | 4 | 14% |
| Over 36 months | 3 | 11% |
The median report in the sample is 18.7 months old. Three of the seven sit between 24 and 26 months, so the threshold does real work: on a 36-month rule the count would be three rather than seven. Publish your threshold in the solicitation and the argument disappears.
One report in the 28 has no Report Date element at all. It supplies a “Completion Date” and a “Revision Date” under a heading structure that is not the template’s. The template sets the floor plainly. Its Essential Requirements list the Report Date element as the “Date of report publication”, with the instruction “At a minimum, provide the month and year of the report publication.” That report is also the one naming a non-ITI edition, and it is the library discovery website, June 2024.
Age is not the same as the rubric’s flag 3, which asks whether the report date and product version match the release being offered. That cannot be coded from a document alone, because it needs the solicitation and the offered build. Age is the part a reviewer can code off the cover page, and it is the part that tells you whether to ask.
Result 3: the required test-method field is populated and still does not answer the question
“Evaluation Methods Used” is one of the template’s minimum content elements, and the instruction attached to it in the Essential Requirements is to “Include a description of evaluation methods used to complete the VPAT for the product under test.” On presence, the sample is close to clean: 27 of 28 carry the element under the template’s own heading, and the 28th carries comparable content under the heading “Evaluation Methods”.
Depth is where the sample separates. Measured strictly inside that element, and counting only the words inside it:
| Words inside the Evaluation Methods Used element | Reports |
|---|---|
| No element under the template’s heading | 1 |
| 1 to 14 words | 10 |
| 15 to 100 words | 10 |
| More than 100 words | 7 |
The median is 37 words. The shortest populated element in the sample is three words long: an automated tool name and its version number, with no method, no scope and no assistive technology. The longest runs 387 words and names the standard and levels, two operating systems with their browsers, the tools and two screen readers. That spread is the finding. The same required field, answered by two vendors in the same market, can be a test plan or a product name.

View the data as a table
| Shortest populated element | Longest element | |
|---|---|---|
| Words inside the element | 3 | 387 |
| What it names | An automated tool name and its version number | The standard and levels, two operating systems with their browsers, the tools |
| Assistive technology | None named | Two screen readers |
| What the reviewer gets | A product name | A test plan |
Two of the 28 populate the required element with a statement of tester familiarity rather than a method. One reads, in full: “Testing based on general product knowledge.” The other reads “Evaluation is based on general product knowledge and testing of website.” Both are answering the first item on ITI’s list of things an author may enter under the element, which reads “Indicate whether testing is performed by testers with knowledge of general product functionality”, followed by ITI’s own instructional note that this means the tester knows the common uses and flows of the product in addition to accessibility. It is a disclosure about who tested, not a description of what was done. ITI evidently found the phrase confusing too: the change-tracking file records that revision 2.4Rev exists in part to “clearly describe what is meant by ‘Testing is based on general product knowledge’”.
That leaves rubric flag 7 with two prongs to score, absent or non-responsive, and they do not give the same answer.
On the absent prong, flag 7 fires on none of the 28. On the non-responsive prong it is arguable on those two, because the element’s own instruction heading in Best Practices for Authors is “Describe the testing performed”, and neither sentence describes any testing. This study scores both prongs clean, for a stated reason: the published rubric’s gloss on flag 7 is to score the field and not its richness, and a vendor who wrote one of those sentences will reply that it entered the first item ITI lists. That reply survives a debrief. A reviewer who reads “non-responsive” more strictly should score those two reports at zero on flag 7 and let the published evidence cap do the rest, since flag 5, 6 or 7 at zero holds the total at 64 whatever the other eleven flags produce. Either way the useful move is the same, and it is not an argument about the template.
One publisher’s five reports each populate the element with the same ten or eleven word list: product knowledge, a monitoring platform, a screen reader, a browser-based checker and browser developer tools. Five of the 28 counts come from one house style, which is exactly why the publisher concentration is disclosed above.
If you want the version of this question you can actually put to an offeror, the wording is in will you accept our test evidence.
Result 4: assistive technology named, versions absent
| Coded field | Reports | Share |
|---|---|---|
| Names at least one assistive technology anywhere in the report | 21 | 75% |
| Names at least one specific tool or assistive technology inside the Evaluation Methods Used element | 23 | 82% |
| Names no specific tool or assistive technology inside that element | 5 | 18% |
| Names an assistive technology together with its version | 3 | 11% |
Three of 28 is the single most reviewable gap in the sample, and it is not a template violation. ITI marks it optional, in the Best Practices for Authors section of the template file itself, in these words: “Describe testing conducted with assistive technologies (Optional: Include the assistive technologies that were used in testing.)” A vendor that omits assistive technology versions has broken nothing. The demand has to come from your solicitation, and citing the template for it is how a reviewer loses an argument with a capture manager.
What the three reports that do it look like: one names three annual releases of a commercial screen reader plus a version of a free one; one names a screen reader build and the browser build it was driven in; one names a screen reader release from 2021 and another from 2022, which is itself informative, because it dates the testing rather than the report. A fourth report gives a tool version and no assistive technology version, which is the three-word element described above.
Screen reader and browser versions are the difference between a claim you can reproduce and a claim you can only believe. Which combinations are worth naming in a solicitation is covered in assistive technology test targets.
Result 5: no Supplemental Accessibility Report anywhere in the sample
Zero of the 28 reports supply a Supplemental Accessibility Report. The word “Supplemental” appears in five of the 28, and every one of those five was read in context: in each case it is ordinary prose about supplemental content, never a report heading.
This one points back at the buyer. The SAR is where the buy-side guidance puts evaluation methods, features that help achieve accessibility, core functions that cannot be used by persons with disabilities, and configuration and installation guidance. Section508.gov’s buy-side guidance is also where the pressure point sits, because it states that “To be considered for award, the ACR must be complete, and submitted according to the instructions.” If your solicitation never asked for a SAR, its absence across 28 reports is a gap in solicitations rather than a gap in vendor practice, and the fix is a clause rather than a rejection. The clause library is in Section 508 contract clauses and the QASP.
The rubric’s flag 10 carries a scoring rule worth restating here: score the content, not the heading. Partial credit depends on how many of the SAR’s four elements survive inside the ACR itself, and this study coded only one of them, the evaluation-methods element, which is present in 27 of 28. So read the zero as “no report supplied the artifact”, not as “flag 10 fires at full weight 28 times”.
Result 6: Supports rows that do not survive their own remarks
This is the one judgment-dependent metric, so the rule comes before the count.
ITI’s definition, which the rubric turns on: “Supports: The functionality of the product has at least one method that meets the criterion without known defects or meets with equivalent facilitation.” The published fire rule, quoted from the rubric: fire only “when no conforming method survives the remark: a defect in the only method available, or a fix scheduled for a future release.” A defect in one of several methods does not fire it. A documented workaround alongside a path that still conforms does not fire it. An exception that the success criterion itself allows does not fire it.
How the metric was produced, in two passes. First, a keyword detector surfaced every Supports row whose remark contained one of a fixed list of defect markers: exception, known issue, known defect, known limitation, will be fixed, will be addressed, will be remediated, will be resolved, will be corrected, future release, next release, upcoming release, roadmap, scheduled, workaround, filed as bug, internal ticket, planned. Second, every detector hit was read in full against the fire rule. The detector over-fired heavily, and no row below was confirmed on a detector hit alone.
Confirmed in 5 of the 28 reports, across 7 rows. Report this as a floor, not an estimate: the detector is keyword driven, so a Supports row describing a defect without any of those markers would have been missed, and one coder applied the rule rather than two.

View the data as a table
| Do | Don’t |
|---|---|
| Fire it only when no conforming method survives the remark: a defect in the only method available, or a fix scheduled for a future release. | Fire it on a defect in one of several methods, or on a workaround that sits alongside a path which still conforms. |
| Publish the fire rule before the count, since this is the one judgment-dependent metric. | Fire it on an exception the success criterion itself allows, such as confirming passwords at 3.3.7 Redundant Entry. |
| Read every detector hit in full against the rule, because a keyword detector over-fires heavily. | Confirm a row on a keyword alone: “no known defects” matches the detector on the word defect inside a negation. |
| Report the confirmed count as a floor rather than an estimate: 5 of 28 reports, across 7 rows. | Rescore a report’s own declared vocabulary: its 31 “Supports with Exceptions” rows carry ITI’s Partially Supports definition. |
| Report (category, report month) | Criterion | Level as stated | The remark that contradicts it |
|---|---|---|---|
| Online reference and eBook platform, December 2023 | 2.4.6 Headings and Labels (AA) | Supports | Under a heading “Current Exceptions:” a single bullet, “Some labels may be incorrectly applied or missing.” |
| Online reference and eBook platform, December 2023 | 3.1.1 Language of Page (A) | Supports | The remark opens “Supports with the current possible exception:” and then reports that with a third-party translation plugin in use “the lang attribute does not switch to the new translated language” |
| Ethnographic research database, February 2026 | 2.1.1 Keyboard (A) | Supports | ”An exception is noted for the Filters. This is due to a 3rd party library. See conformance and remediation documentation for details on a fix.” |
| Library service application B, September 2024 | 2.4.6 Headings and Labels (AA) | Web: Supports, Authoring Tool: Partially Supports | ”Web: There are some small exceptions that were filed as bugs that we will need to address”, followed by two numbered items naming modal titles set as H4 on pages with no H3, and results information nested in an H2 |
| Library service application B, September 2024 | 2.4.1 Bypass Blocks (A) | Web, Electronic Docs and Authoring Tool: Supports | ”Authoring Tool: A small exception is that there is no way right now to keyboard bypass the user list names on the Appointments -> Calendar interface to get to the main calendar content.” |
| Library service application C, October 2024 | 1.4.11 Non-text Contrast (AA) | Web: Supports, Authoring Tool: Partially Supports | ”Web: The Cancel button background on the blog comment section is similar to the blog background color. Users can still find the Cancel button by its context.” |
| Library service application D, December 2024 | 3.2.3 Consistent Navigation (AA) | Web, Electronic Docs and Authoring Tool: Supports | ”Authoring Tool: There is an exception with the widget preview screens. The widget preview screens do not have the common top nav bar.” |
The 1.4.11 row is the weakest of the seven and the one a vendor would contest first, so it is printed in full rather than trimmed. The vendor’s second sentence, that users can still find the Cancel button by its context, is a statement about impact, not a second method. Success criterion 1.4.11 is a contrast threshold, and locating a control by context does not meet a contrast threshold or supply equivalent facilitation for it; the remark records a known defect in the only Cancel button on that screen. That is the reading applied here. A reviewer who accepts the vendor’s sentence as sufficient should record 6 rows in 4 reports instead, and either way the action is a question to the vendor rather than a rejection. Note also that the same cell ends with “An internal ticket has been filed for these issues”, which covers the Authoring Tool defect in the same cell, so it is not quoted above as though it attached to the Web row.
What the rule kept out matters as much as what it let in, because these are the rejections that stop the metric from being an opinion.
- Remarks reading “no known defects” were matched by the detector on the word “defect” inside a negation. Rejected.
- Rows at 3.3.7 Redundant Entry citing “the allowed exception of confirming passwords during account creation” invoke the success criterion’s own exception. Rejected.
- Rows at 1.4.5 Images of Text citing graphs and equations that must be preserved for meaning, and rows at 1.1.1 citing purely decorative images, invoke exceptions the criteria themselves carry. Rejected.
- One report uses “Supports with Exceptions” where the template has “Partially Supports”, declares the substitution in its own Terms section, and gives it ITI’s Partially Supports definition verbatim: “Some functionality of the product does not meet the criterion.” The string appears 32 times in the file, one of which is that Terms definition, so 31 are row-level uses, and “Partially Supports” appears nowhere. Counting those 31 rows as contradicted Supports rows would have taken the headline from 7 rows to 38. Under the report’s own declared definitions the rows are internally consistent, so all 31 were rejected here. This is the sample’s biggest false-positive trap.
- Two further detector hits in one report were read and not confirmed, and one row asserting that an exception is “valid” without naming a basis was left unconfirmed. It is a question for the vendor, not a finding.
The 31-row case is worth one more sentence, because it is not a vendor inventing vocabulary. “Supports with Exceptions” is ITI’s own retired level. ITI’s VPAT page records that “in a previous update of the VPAT, ‘partially supports’ replaced ‘supports with exceptions.’ This change was made at the request of representatives of the U.S. Access Board”, and the change-tracking file puts that swap in version 2.2, June 2018, for the 508 and INT editions. The report claims VPAT 2.4 on its cover, names no edition, and runs pre-2.2 vocabulary throughout: it also prints “2017 Section 508” 43 times, which is the label 2.2 replaced with “Revised Section 508”. The correct finding is not a coinage. It is a report whose Terms section is three revisions behind the version it claims, which is a template deviation and is recorded as one below.
One report is recorded separately, because the defect is remark specificity rather than a level contradiction: under three Supports levels at 1.3.1 Info and Relationships it announces that a monitoring tool “has found some minor exceptions”, follows the colon with no list at all, and leaves a second scope label in the same cell with no text after it. That is the rubric’s flag 6 territory, and flag 6 is not reported across this sample. Row-level extraction from multi-column PDF tables was not reliable enough to publish: four reports left between 9 and 55 rows with no parseable conformance level, which is almost certainly an extraction artifact rather than a missing level. Publishing a remark-specificity percentage off that would have been an opinion dressed as data.
Result 7: defects visible without reading a single conformance row
Six of the 28 reports carry at least one of the defects below. Five of the six are template deviations, which is rubric flag 12, and the sixth is the “Not Evaluated” case, which is flag 4. The list of defect types was fixed by reading rather than by an exhaustive template audit, so treat both counts as floors.
| Deviation | Reports | What a reviewer sees |
|---|---|---|
| Blank-template front matter still attached | 3 | ITI’s instructions to vendors sit ahead of the report, so the buyer receives the manual and the report in one file. Two of the three also show unresolved Microsoft Word field codes where a table of contents should be. A reader searching for “Report Date” hits the instruction text first. |
| Conformance level renamed | 1 | ”Supports with Exceptions”, the level ITI retired in 2018, on 31 rows, declared in the report’s own Terms. |
| No Report Date element, and a non-ITI edition name | 1 | ”Completion Date” and “Revision Date” in place of the template’s Report Date. |
| Product Description describes a different product | 1 | The description in one report is the same sentence used in another report from the same publisher, and it describes the other product. |
| ”Not Evaluated” on Level A or AA rows | 1 | Six rows, all of them WCAG 2.2-only criteria. |
The front matter row and its mirror image are worth reading together, because one of them is not a defect at all. ITI’s instruction to vendors is unambiguous: “When publishing your Accessibility Conformance Report, be sure to remove the entire first 10 pages of this document, including the table of contents, introductory information and instructions.” Section508.gov’s sell-side guidance says the same thing and names the heading to cut to. So three of these 28 shipped pages their vendor was told to delete. A fourth report deleted them correctly and looks worse for it: its page-number field still counts from the untrimmed template, so a complete seven-page ACR opens at “Page 10 of 16” and repeats that pagination to the end. It carries the full heading block, contact information with an email and a phone number, Evaluation Methods Used, Terms, and both the Level A and Level AA tables. Nothing is missing. It is recorded here as a reading trap rather than a deviation, because a reviewer who rejects a report for starting at page 10 will be rejecting the vendor who followed the instructions.
Two of the table rows need their reading stated, or the count gets used to say something it does not support.
The copy-forward Product Description is the cleanest evidence in the sample that a finished document went out unread. It is a documentation defect and nothing more. It says nothing about either product.
The “Not Evaluated” case is a template-form defect, not a vendor declining to test what it claimed. The six rows are 3.2.6 Consistent Help and 3.3.7 Redundant Entry at Level A, and 2.4.11 Focus Not Obscured (Minimum), 2.5.7 Dragging Movements, 2.5.8 Target Size (Minimum) and 3.3.8 Accessible Authentication (Minimum) at Level AA, all of them WCAG 2.2-only criteria, and the same report’s Applicable Standards table declares WCAG 2.2 Level A and AA out of scope. The template restricts the level plainly: “Not Evaluated: The product has not been evaluated against the criterion. This can only be used in WCAG Level AAA criteria.” Either “Not Applicable” or an omitted 2.2 table would have said the same thing in a template-consistent way. Which WCAG version a given US rule actually requires is a separate question, set out in WCAG version requirements by rule.
Flag 12 has a consequence that reviewers underuse. The template’s Essential Requirements for Authors state that “Users of the VPAT agree not to deviate from the Essential Requirements for Authors”, and that using the template and the name requires the registered service mark. A report that deviates materially is not just untidy, it is a document whose author has stepped outside the terms under which the template may be called a VPAT at all. That is a reason to ask for a reissue, and it is a small ask.
Contact information, coded on the template’s own bar that “Listing an email is sufficient”: one report of 28 leaves the Contact Information element empty, with no email address anywhere in its 55 pages. Two more give no reachable route, both reading that the reader should contact an institutional sales representative. That is flag 11 firing once and arguable twice.
Which of the twelve flags a document study can score
The rubric’s twelve flags and weights are reproduced here unchanged, with what this study could and could not code against each.
| # | Flag | Weight | Coded here | Result across the 28 |
|---|---|---|---|---|
| 1 | Applicable Standards/Guidelines indication absent, or inconsistent with the edition used | 8 | Partly | Not coded across the sample. Two instances found by reading: one report whose 2.2 rows read “Not Evaluated” while its own scope table declares 2.2 out of scope, one report naming a non-ITI edition. |
| 2 | Federal 508 obligations not reported (a WCAG-only claim against a 508 buy) | 10 | Yes | Fires on 5 of 28, the five WCAG-edition reports with no 302.x and no 602.x table. |
| 3 | Report date or product version does not match the release being offered | 8 | No | Needs the solicitation and the offered build. Age coded instead: 7 of 28 over 24 months. |
| 4 | ”Not Evaluated” appearing on a Level A or Level AA success criterion | 8 | Yes | Fires on 1 of 28, 6 rows, all WCAG 2.2-only criteria. |
| 5 | A “Supports” row whose own remark leaves no conforming method standing | 12 | Yes, under the published fire rule | Confirmed in 5 of 28, 7 rows. A floor. |
| 6 | Thin or missing remarks on “Partially Supports” and “Does Not Support” rows | 12 | No | Row-level extraction was not reliable enough to publish. Not reported. |
| 7 | ”Evaluation Methods Used” absent or non-responsive | 12 | Yes | Fires on 0 of 28 on the absent prong. Arguable on 2 of 28 on the non-responsive prong, scored clean here for the reason given in Result 3. |
| 8 | Assistive technologies and testing tools not named | 4 | Yes | 5 of 28 name none inside the element. 3 of 28 name an assistive technology with a version. |
| 9 | Chapter 3 Functional Performance Criteria not answered where Chapters 4 and 5 leave a function uncovered | 10 | Conditional, not coded | Firing it requires establishing which product functions the hardware and software chapters do not reach. Table presence coded instead: 23 of 28 carry a 302.x table. |
| 10 | No Supplemental Accessibility Report | 8 | Yes | 0 of 28 supply one. |
| 11 | No contact information for follow-up questions | 4 | Yes | Empty element in 1 of 28. No reachable route in 2 more. |
| 12 | Material deviation from the template’s essential requirements | 4 | Yes, for a fixed list of deviation types | At least one recorded deviation in 5 of 28. |
No total scores and no score distribution are published, and the reason is a rule rather than caution. Flag 3 needs the offered release. Flag 6 needs reliable row-level extraction this sample did not have. Flag 9 is conditional on a product-function finding no document study makes. The rubric also caps any total at 64 when flag 5, 6 or 7 scores zero, so a total assembled from nine flags out of twelve would not be the instrument’s output. It would be a different number wearing the instrument’s name.
That is the honest answer to the question the sample was built to ask. Of twelve flags, seven can be coded from the document on your screen and five cannot. Seven is enough to tell you whether to accept, return, or spend money on testing.
What the document does not tell you
It does not tell you whether the product conforms. A report can be current, on the right edition, specific about method, name assistive technology versions, and still describe a product with defects nobody found. The reverse holds too: a thin report can describe a well-built product whose vendor writes badly. Documentation quality is a proxy for the diligence behind a claim, not a measurement of the claim.
It also does not tell you what a script or widget bolted onto the product achieves, since that claim lives in the product and not in the document.
There is no central federal source to check a vendor’s report against. GSA’s own ACR Library states that reports “are provided only for the tools and training made available by GSA through Section508.gov”. Counted on July 27, 2026 it holds nine rows, three tools and six online training courses, three of them showing a Report Date of “Pending” with no downloadable report. It contains no third-party vendor ACRs. You evaluate what the offeror hands you.
And the verification gap on the buyer’s side is documented, though it has to be quoted precisely. GSA’s FY 2025 governmentwide Section 508 assessment records agency self-report on how often ICT deliverables from a contract are verified for Section 508 conformance: never 7 percent, rarely 20 percent, sometimes 23 percent, often 13 percent, almost always 30 percent, and 7 percent recording that they do not perform the activity. That is a statistic about agency acquisition practice, self-reported, and the assessment states its own limitation: “Agencies, parent agencies, and components self-reported the data in this report. No independent validation or external data was utilized.” It is not a measure of ACR accuracy and it does not mean 30 percent of ACRs go unverified.
What to put in the next solicitation
The three fields this study found thinnest are optional in the template, which means they arrive only if you ask for them in writing. All three cost a vendor a paragraph.
- Assistive technologies and tools, with versions. Optional in ITI’s Best Practices, absent with versions in 25 of the 28 reports here. Ask for the assistive technology, its version, the browser and its version, and the operating system. Ask through the Supplemental Accessibility Report, not as a template requirement.
- A Supplemental Accessibility Report for each standard commercial item. Zero of 28 supplied one. This is the artifact that carries evaluation methods, features that help achieve accessibility, core functions that cannot be used by persons with disabilities, and configuration and installation guidance, and it is the correct home for the two asks above and below.
- A test-method description that names scope. The required element is populated in 27 of 28 and reaches 14 words or fewer in 10 of them. Specify what “described” means: which platforms, how many screens or documents, which standard and level, manual as well as automated, and who performed it.

View the data as a list
- Assistive technologies and tools, with versions: Optional in ITI’s Best Practices, absent with versions in 25 of the 28. Ask for the version, the browser and the operating system.
- A Supplemental Accessibility Report: Zero of 28 supplied one. It is the correct home for evaluation methods, unusable core functions and configuration guidance.
- A test-method description that names scope: Populated in 27 of 28, and 14 words or fewer in 10 of them. Specify platforms, screens, standard and level, manual as well as automated.
Then state your own thresholds. Publish the maximum report age you will accept, publish that the edition must be Revised 508 or INT for a federal buy, and publish that completeness is measured against your ACR instructions. Section508.gov’s buy-side guidance already gives you the award condition; your instructions are what make it operable.
The sentence that goes in the file is not “this ACR looked weak”. It is closer to this: this firm applied a published twelve-flag instrument consistently across offerors; against a 28-report sample of published ACRs scored on the same instrument, this report is past the 24-month age threshold that 21 of the 28 met, states its evaluation method in fewer words than the 37-word sample median, and names no assistive technology version, a field 3 of the 28 populate; we are therefore exercising the reserved right to test before award. That is a defensible record whichever way the decision goes. Where several offerors land in the same place on the same field, that is a market finding rather than a vendor finding, and it belongs in the file as one.
Your next step
Take the ACR on top of your pile and code five fields against the base rates above. Template edition, as named on the report. Report date, and its age today. The Evaluation Methods Used element, word for word, with a word count. Whether any assistive technology carries a version. Whether a Supplemental Accessibility Report exists. If the report is over 24 months old, states its method in 14 words or fewer, and names no assistive technology version, it sits with the 7 of 28, the 10 of 28 and the 25 of 28 respectively. Five fields, three counts, and a finding you can write down instead of an impression you have to defend.
If the gap turns out to be on your side of the table, ADACP’s Section 508 procurement support covers ACR review and the solicitation language that produces reviewable reports in the first place, including the Supplemental Accessibility Report clause that every finding above depends on. Vendors reading this from the other side: vendor guidance on ACRs and VPAT testing apply the same twelve flags before a report is published, which is the cheaper end of this exercise.