Coverage is broad, not complete
536 filings from 96 organizations is a lot of the market and not all of it. Every document that was looked for and not captured is recorded — 490 of them — because a dataset that quietly drops what it missed is asking to be trusted rather than checked.
What is actually blocked
127 documents are confirmed to exist and could not be read by a machine. About a third are pages that load and carry no metrics report where one was promised. Most of the rest are obstacles on the payer's side: 33 paths disallowed by robots.txt, 16 that answer every request with a refusal, a bot-protection challenge or two, and a few published as scanned images or spreadsheet binaries rather than text. Humana's assembles in the browser and has no underlying document at all. A further 5 are ours rather than theirs — long consolidated PDFs read only in part, whose captured contracts are in the dataset.
This is the group worth arguing about. CMS required these filings to be published. Publishing one in a form no machine can read satisfies the letter of that and not much else.
What was simply not found
The larger number, 304, is honest failure on this end rather than obstruction on theirs. Most are addresses guessed from a pattern that worked for a sibling brand and did not generalize; the rest are payers in scope whose filing was never located. Some of those documents almost certainly exist somewhere this crawl did not reach.
Keeping the two apart matters. Reporting all 490 as withheld documents would overstate the finding by roughly three times, which is exactly the sort of arithmetic this project exists to catch.
Every gap is listed individually, with its reason and its classification, in out/coverage_gaps.csv — see the pipeline repository. The classification is assigned in pipeline/gaps.py and stored alongside the data, so this page, the CSV and the database always agree.