φ(ai) PHI AI
DocumentationDocs
Repository

Fairness screen

Section 1557 / 92.210, run for real: the Denial risk model (registry #11) examined across sex, age and payer subgroups of the live corpus — the model's mean score against each group's actual denial rate, on 667 sampled claims. The gap is the finding; the thresholds are printed beside it. Aggregates only — no patient appears on this plane.

Subgroupn Actual denial % Mean model score GapVerdict (≤5 calibrated · ≤12 watch · else review)
age · 45-64 171 9.9% 9 0.9 calibrated
age · 18-44 228 7.9% 8.9 1 calibrated
age · 0-17 143 10.5% 9 1.5 calibrated
age · 65+ 125 11.2% 9 2.2 calibrated
payer · Medicare 106 9.4% 9 0.4 calibrated
payer · Medicaid 113 8% 9.1 1.1 calibrated
payer · Blue Cross Blue Shield 106 10.4% 9 1.4 calibrated
payer · Humana 46 4.3% 8.1 3.8 calibrated
payer · UnitedHealthcare 120 13.3% 9.1 4.2 calibrated
payer · Aetna 88 13.6% 9 4.6 calibrated
sex · female 346 9.2% 9 0.2 calibrated
sex · female 346 9.2% 9 0.2 calibrated
sex · male 321 10% 9 1 calibrated

Subgroups under 25 sampled claims are computed but not published here — small cells identify people. Screen computed 2026-08-29 23:30:25, cached one hour; the preflight gate (6.2) reads this screen's existence for the claims slot.

✦ AI Reading the screen

Fairness review drafted by claude-sonnet-5 · 2026-08-29 20:16:02 · informs governance, never adjudicates
READING THE SCREEN
All twelve subgroup rows are tagged "calibrated," but the underlying gap between actual denial rate and mean model score is not uniform. Among age bands (0-17, 18-44, 45-64, 65+), gaps run from 0.9 to 2.2, rising with age. Among payers, gaps span a much wider range: Medicare 0.4, Medicaid 1.1, Blue Cross Blue Shield 1.4, Humana 3.8, UnitedHealthcare 4.2, Aetna 4.6. Sex shows the tightest gaps (female 0.2, male 1.0), but the female row (n=346, 9.2%, score 9, gap 0.2) is listed twice, identically — that looks like a duplicate record rather than two distinct subgroups, and should be corrected before any conclusions are drawn from the sex comparison.
WHERE TO LOOK CLOSER
1. Payer subgroup, not age or sex, shows the largest calibration gaps. UnitedHealthcare (n=120, actual 13.3%, gap 4.2) and Aetna (n=88, actual 13.6%, gap 4.6) both have the highest actual denial rates in the table and the largest gaps — the model is under-scoring relative to observed outcomes for these two payers. Humana (n=46, actual 4.3%, gap 3.8) is the opposite pattern: lowest actual denial rate but still a large gap, and it also has the smallest sample size of any row, which makes its gap the least statistically stable.
2. Age 65+ (n=125, gap 2.2) has the largest gap in the age category, though it is modest compared to the payer gaps and the sample size is also the smallest among age bands.
3. The "calibrated" label appears to be applied uniformly regardless of gap size (0.2 through 4.6 all pass), which suggests the calibration threshold used to generate this flag may be too permissive to catch the payer-level spread described above.

SPEC 6.3, live: a deterministic screen over a deterministic model, a language model reading the results out loud, and a paper trail either way. Synthetic data throughout.