Which hospitals charge more than their services explain
Every price in the sample is two things at once: the procedure (an MRI costs more than a blood draw everywhere) and the hospital (some places charge more for everything). A model over 12.1M published negotiated rates splits the two apart. What is left, per hospital, is its fair-price index: ×1.20 means its rates run 20% above what its own mix of services would predict at the average hospital. None of this is cause, and none of it is a judgment about care — it is the structure of the prices hospitals themselves published.
Every scored file, on one axisⓘ
orange = at least double, or at most half, its expected prices · click a dot for the file's numbersWhat this measures. Hospitals now have to publish the rates they negotiate with insurers. We read 12.1M of them across 1,202 hospital price files, and asked one question of each hospital: given what it sells, does it charge more or less than everyone else selling the same things? The typical item in the panel goes for about $612; a hospital at ×2 charges roughly double the going rate for its own mix, across the board.
What we found. Half of all files sit between ×0.76 and ×1.30 of their expected prices — but the tails are long: 74 files charge at least double what their mix predicts, and 51 charge half or less. The spread runs from ×0.25 to ×4.09 — the same basket of care, priced 16 times apart depending on the door you walk through. Together, "what the item is" plus "which hospital sells it" explain 91% of all price variation — being at an expensive hospital is not a procedure-by-procedure accident; it is the hospital.
The unit is the published file, not the building. A system that publishes one file per campus appears once per campus; rows carry the CCN so files for one hospital can be grouped. The negotiated median summarises whatever payers the file listed — payer identity is not in the data, so payer mix and pricing posture cannot be told apart.
Where the markups live
The index, state by state: each state is shaded by the median index of its scored files, so the map shows where hospitals as a group price above or below what their mix of services predicts. Under it, the files at both ends of the national ranking, each with its interval.
The twelve most expensive
index · 95% interval · click for the numbersThe cheapest eight
same scale — the orange rule is ×1The pattern in the extremes is not random: specialty and children's hospitals crowd the top, public safety-net and critical-access hospitals the bottom. The interval is wide where a file prices few items — the count is printed on every row. The unit is the published file, not the building. A system that publishes one file per campus appears once per campus; rows carry the CCN so files for one hospital can be grouped. The negotiated median summarises whatever payers the file listed — payer identity is not in the data, so payer mix and pricing posture cannot be told apart.
Can you tell an expensive hospital without opening its file?
Partly. The same trees-plus-ridge pair as the county models below, aimed at the index itself: from a hospital's size, its system, and its market, the model explains about 30% of who charges what (cross-validated R² 0.30 ± 0.03). The rest — the distance from the diagonal in the chart — is a premium or a discount that nothing we can list about the hospital explains.
What the model relied onⓘ
permutation importance · ridge coefficientThe strongest signals are how much a file prices (broad, everything-priced files skew cheaper — partly a disclosure effect), how big the hospital's system is, and the market it sits in: counties with more uninsured residents and higher incomes carry dearer hospitals. But the trees beat the straight line by a wide margin (ridge R² 0.03 against 0.29 held out) — pricing power is not a tidy linear story, it is thresholds and combinations. And 70% of the variation stays unexplained by anything on the list: two hospitals of the same size, in the same kind of system and market, can still price a world apart.
Ownership is deliberately absent: the AHRQ ownership codes arrive without their labels for a third of hospitals, and a number nobody can read has no place on a page. State is absent because fifty dummies would soak up the market patterns the listed features are here to explain. Nothing here establishes cause.
Expected vs. actualⓘ
click a dot for the fileThe prices hospitals didn't publish, estimated
No hospital prices everything: the hospital-by-item table is 28% empty even for the items in this panel. On top of the item and hospital effects, a low-rank model learns who deviates on what — the signature of a hospital that discounts imaging but not surgery — and fills the blanks. We grade it only on prices it was never shown.
How wrong is it?ⓘ
held-out cells: 547,456Predicting a missing price as "the going rate for the item, adjusted by the hospital's index" is already decent — off by 38% on a typical held-out price. Letting the model learn 40 patterns of who-deviates-on-what cuts the typical miss to 11% (held-out R² 0.979). The bands on every estimate come from those same held-out errors: 80% of the time, the real price fell inside the band.
Try it: one item, one state
click a row for that file's index and estimatesEstimates are model output, not published prices — never quote one as a hospital's rate. They exist to make the published numbers comparable: a hospital that skipped an item no longer vanishes from the comparison, it appears with an honest band around what its overall pricing suggests.
What travels with medical debt and early death
From hospitals to the places around them. For each outcome below, a model learns to predict a county's value from 17 things we know about the county — what its hospitals bill, who has insurance, what people earn and whether they have enough to eat, who lives there and how rural it is, how far the nearest doctor is and how many smoke. Then we ask the model which of those it relied on, and how its prediction moves as each one changes. None of this is cause. It is the structure the data has; every figure carries a note on how it was made.
What this measures. The share of adults in a county who have a medical bill that went unpaid long enough to be sent to a collections agency and land on their credit report. The Urban Institute measures it from a 4% sample of credit records (August 2025). Seven states keep medical debt off credit reports entirely, so their counties are not in this model. The model sees 2,654 counties and, for each, 17 facts about the place — what its hospitals bill, who has insurance, what people earn and whether they have enough to eat, who lives there and how rural it is, how far the nearest doctor is and how many smoke. The hospital billing figures exist only for the 1,181 counties with a hospital in the Medicare files; for the rest the model is told there is no figure.
How well it does. Shown counties it had never seen, the model explained a little under half of the difference between them (44% of the variation). Its typical miss is 2.2 percentage points, against a county average of 5.2%. That leaves a lot unexplained: things this data does not contain — local hospital billing practices, state law, who sues whom — matter at least as much as what it does.
What it leaned on. Most of all, uninsured share (under 65) — where that is higher, the share of adults with medical debt in collections is higher. Then Non-Hispanic Black share — where that is higher, the share of adults with medical debt in collections is higher. Then food-insecure share — where that is higher, the share of adults with medical debt in collections is higher. What hospitals bill per dollar paid ranked no higher than 12 of 17: once uninsured share (under 65) and Non-Hispanic Black share are known, the model barely used the price data. High list prices are where a bill starts; whether anyone is left holding it is about coverage.
Read "leaned on" as "found useful for guessing", not "caused". Two things that move together across counties — say, smoking and poverty — share the credit, and the model cannot tell them apart. Every number here is read from the model tables; the sentences follow the numbers.
What the model relied onⓘ
permutation importance · ridge coefficientⓘ · click a bar for the numbersLimits. One year, cross-sectional, county level. Nothing here establishes cause. Seven states bar medical debt from credit reports and are absent from the debt model. CDC PLACES disease rates are not modelled as outcomes: they are themselves estimated from the same demographic covariates, so a model of them would mostly recover CDC's model. PLACES smoking, obesity and checkup rates are used as predictors, with that caveat; its uninsured rate is not, because the ACS figure measures the same thing directly. Medicare price ratios exist only for counties with a hospital in the Medicare files (43%); elsewhere the trees see a gap. Age and poverty are absent from this build (the ACS source was not fetched).
How the prediction movesⓘ
predicted Share of adults with medical debt in collections, hold-out counties, 2nd–98th percentile of each feature · click a curve for its grid points+3.9% across range
+1.3% across range
+1.1% across range
+1.1% across range
+1.0% across range
−0.6% across range
−1.4% across range
+0.5% across range
All panels share one vertical scale (3.3% – 7.9%), so a flat line is flat and a steep one is steep in the same units. Features with no curve were not among the eight the model relied on most.
Every county, shaded by each outcome
The same outcomes the models predict, on the ground. The switches under the map change the outcome; the selectors on the right zoom to one state and pull up one county's numbers. Clicking any county opens the data behind it.
The model cards
The hospital index and the estimates
The county models
Code and every parameter: analysis/hospital_models.py and analysis/county_models.py in the app repository. Rerunning either with the same seed (20260821) reproduces every number on this page. The model tables are in the API (set=hospital_model) and set=model.
Sources
the files read, as fetched — not our copy of themThe county models read the county profile the pipeline builds from these files; the hospital models read the sampled price files themselves (every file's URL, fetch time and SHA-256 is on its own page via the API). analysis/hospital_models.py and analysis/county_models.py record the seed and every parameter.
- Debt in America 2025 — county-level medical debt · Urban Institute · August 2025 credit-bureau panelDebt%20in%20America%20County-Level%20Medical%20Debt.xlsx · 396K bytes · fetched 2026-08-21 · sha256 4b2ec3982817… · publisher's page · ODC-BY
- County Health Rankings 2025 — national analytic data · University of Wisconsin Population Health Institute · 2025 release (measures 2019–2023)analytic_data2025_v3.csv · 13.1M bytes · fetched 2026-08-21 · sha256 5a129b856489… · publisher's page · free for non-commercial use with attribution; see countyhealthrankings.org
- CDC PLACES — county health measures, 2025 release · CDC · 2025 release (BRFSS 2022–2023 model-based estimates)rows.csv · 4.8M bytes · fetched 2026-08-21 · sha256 a47cad3a852c… · publisher's page · public domain (US federal)
- ACS 5-year 2023 — county age and poverty · US Census Bureau · 2019–2023
- MUP_INP_RY26_P03_V10_DY24_PrvSvc.CSV · 38.0M bytes · fetched 2026-08-21 · sha256 2ab6da15be4c… · publisher's page · public domain (US federal)
- MUP_OUT_RY26_P04_V10_DY24_Prov_Svc.csv · 28.1M bytes · fetched 2026-08-21 · sha256 f293918edbf6… · publisher's page · public domain (US federal)
- chsp-hospital-linkage-2023.csv · 1.5M bytes · fetched 2026-08-21 · sha256 a86146f10c8d… · publisher's page · public; cite AHRQ Compendium of U.S. Health Systems
- 2020 ZCTA to County relationship file · US Census Bureau · 2020tab20_zcta520_county20_natl.txt · 6.8M bytes · fetched 2026-08-21 · sha256 3ed41278d637… · publisher's page · public domain (US federal)
Every figure on this page is computed from these files by pipeline/prices/; the same rows are in the API.