φ(ai) PHI AI
DocumentationDocs
Repository

Population health & quality measures

A population is described by its structure, its mortality, its burden of disease, and how unevenly those fall — then by what a model can say about who is at risk, and how honestly that model reports itself. Every figure on this page is computed from the Data Store on load: prevalence with exact confidence intervals, age-standardisation against the 2000 US standard population, and a logistic model fitted here, on this corpus, with its discrimination, calibration and subgroup performance all shown.

33,311
people in the population
16,766
with 2+ active conditions — 50.3% multimorbid
476
deceased — 1.43% of the cohort
0.922
held-out AUC of the mortality model

1 · The shape of the population

Age and sex structure — the first thing any population health assessment establishes, because every rate below is either driven by it or has to be adjusted for it.

2 · Mortality by age, and what it implies

Observed deceased proportion in each attained-age band, with 95% Wilson intervals — wide where the band is small, and honestly so. The survival curve beside it is an illustrative period life table: it reads each band's proportion as nqx. The corpus carries no date of death, so this demonstrates the method rather than estimating mortality, and the number below is labelled accordingly. It also explains the implausible e(x) at the top of the table: no one under 35 in this corpus is recorded deceased, so the early nqx are exactly zero and the cohort loses nobody before middle age. A real life table is anchored by infant and child mortality, which synthetic data like this does not carry — read e(x) from age 65 down, where the bands have events in them.

Age bandPopulation DeceasedProportion 95% CI l(x) per 100,000 e(x), illustrative
0-4 1,907 0 0.00% 0.00–0.20 100,000 96.3
5-14 3,638 0 0.00% 0.00–0.11 100,000 91.3
15-24 4,005 2 0.05% 0.01–0.18 100,000 81.3
25-34 4,120 1 0.02% 0.00–0.14 99,950 71.3
35-44 4,234 8 0.19% 0.10–0.37 99,926 61.3
45-54 4,403 16 0.36% 0.22–0.59 99,737 51.5
55-64 4,417 27 0.61% 0.42–0.89 99,375 41.6
65-74 2,684 99 3.69% 3.04–4.47 98,767 31.8
75-84 2,592 167 6.44% 5.56–7.45 95,124 22.9
85+ 1,311 156 11.90% 10.26–13.76 88,995 14.1

3 · Burden of disease — crude and age-standardised

Crude prevalence answers "how many of our people have this." Age-standardised prevalence answers "how would this population compare to another with a different age structure" — the two diverge exactly where age drives the disease, which is the point of computing both. Standardised to the 2000 US standard population; intervals are Wilson (crude) and normal on the weighted sum (standardised).

ConditionPeople Crude95% CI Age-standardisedΔ
Body mass index 30+ - obesity (finding) 7,124 21.39% 20.95–21.83 19.00% -2.39
Essential hypertension (disorder) 6,974 20.94% 20.50–21.38 17.87% -3.07
Hyperlipidemia (disorder) 4,763 14.30% 13.93–14.68 11.86% -2.44
Prediabetes (finding) 4,313 12.95% 12.59–13.31 11.19% -1.76
Allergic rhinitis (disorder) 3,327 9.99% 9.67–10.31 9.33% -0.66
Type 2 diabetes mellitus (disorder) 3,055 9.17% 8.87–9.49 7.81% -1.36
Asthma (disorder) 3,024 9.08% 8.77–9.39 8.48% -0.60
Chronic low back pain (finding) 2,958 8.88% 8.58–9.19 7.78% -1.10
Anemia (disorder) 2,607 7.83% 7.54–8.12 7.03% -0.80
Gastroesophageal reflux disease (disorder) 2,483 7.45% 7.18–7.74 6.54% -0.91
Generalized anxiety disorder (disorder) 2,425 7.28% 7.01–7.56 6.68% -0.60
Major depressive disorder (disorder) 2,167 6.51% 6.25–6.78 5.90% -0.61

4 · Multimorbidity

Counting diseases per person, not people per disease — 50.3% of this population carries two or more active conditions, which is the number that actually predicts complexity of care.

5 · Where the burden falls unevenly

Prevalence by payer with 95% intervals, and the ratio to the lowest group. A ratio is a question, not a verdict: payer tracks age and income, so these gaps are a place to look, not a finding on their own.

Type 2 diabetes mellitus (disorder)

Payern Prevalence95% CI Ratio to lowest
Medicare 5,332 17.16% 16.17–18.20 2.96×
Humana 2,281 11.92% 10.66–13.32 2.06×
UnitedHealthcare 5,563 9.01% 8.28–9.79 1.55×
Blue Cross Blue Shield 5,852 7.33% 6.69–8.03 1.26×
Aetna 4,468 7.21% 6.48–8.00 1.24×
Cigna 3,968 6.98% 6.23–7.82 1.20×
Medicaid 5,847 5.80% 5.23–6.43 1.00×

Essential hypertension (disorder)

Payern Prevalence95% CI Ratio to lowest
Medicare 5,332 37.12% 35.83–38.42 2.58×
Humana 2,281 26.83% 25.05–28.69 1.87×
UnitedHealthcare 5,563 20.31% 19.28–21.39 1.41×
Cigna 3,968 17.54% 16.39–18.75 1.22×
Aetna 4,468 17.30% 16.22–18.44 1.20×
Blue Cross Blue Shield 5,852 16.11% 15.19–17.08 1.12×
Medicaid 5,847 14.38% 13.51–15.31 1.00×

Major depressive disorder (disorder)

Payern Prevalence95% CI Ratio to lowest
Medicare 5,332 9.75% 8.98–10.58 1.93×
Humana 2,281 7.19% 6.20–8.32 1.42×
UnitedHealthcare 5,563 6.79% 6.16–7.49 1.34×
Cigna 3,968 5.75% 5.06–6.51 1.14×
Blue Cross Blue Shield 5,852 5.69% 5.13–6.31 1.12×
Aetna 4,468 5.55% 4.92–6.26 1.10×
Medicaid 5,847 5.06% 4.53–5.65 1.00×

6 · Quality measures

Numerator over denominator, computed live, each with a 95% Wilson interval — a measure reported without one is a point estimate pretending to be a fact.

MeasureDefinitionNum DenRate 95% CI
Controlling high blood pressure Hypertensive patients with a systolic reading under 140 in the last 18 months. 1,875 6,974 26.9% 25.9–27.9
HbA1c testing in diabetes Diabetic patients with an A1c resulted in the last 12 months. 1,141 3,237 35.2% 33.6–36.9
Statin therapy in coronary disease Patients with coronary/ischemic heart disease on an active statin. 412 1,019 40.4% 37.5–43.5
BMI documented (adults) Living adults with at least one BMI on record. 16,760 26,126 64.2% 63.6–64.7

7 · A mortality risk model, fitted here

Logistic regression over 10 features, fitted by iteratively reweighted least squares (Newton-Raphson) with ridge damping — converged in 14 iterations, 2,711 ms. Trained on 26,649 patients, held out on 6,662 by a deterministic split, so these numbers reproduce exactly on every run. 476outcome events in the cohort.

0.922
AUC, held out — training 0.936, so the gap is 0.014
0.0103
Brier score — lower is better calibrated
55 / 79
deaths caught in the flagged top decile
612
false positives at that operating point

Coefficients as odds ratios, with 95% intervals

Ordered by |z|. An interval crossing 1 is a feature the data does not let us claim. These are adjusted associations in a synthetic corpus — not causes, and not transportable to a real population.

Discrimination by subgroup

One AUC over a whole population hides who the model serves badly. Subgroups with at least 200 held-out patients and 5 events are shown; the rest are computed but withheld, because an AUC over four events is noise.

8 · ✦ AI Review of these results

Reviewer's read drafted by claude-sonnet-5 · 2026-08-29 22:45:55 · reads results, never records
What the population looks like
This is a 33,311-person synthetic cohort with 476 deaths (1.43%) and half the population carrying 2+ active conditions. Mortality tracks age steeply and plausibly: near zero through age 34, rising to 0.61% at 55-64, 3.69% at 65-74, and 11.90% at 85+, with non-overlapping CIs across the older bands — this is the expected shape for a mortality outcome, not an anomaly. Crude prevalence is consistently higher than age-standardized prevalence for every condition listed (e.g., hypertension 20.94% crude vs 17.87% ASR, T2DM 9.17% vs 7.81%), confirming this cohort skews older than the 2000 US standard population, so crude rates alone would overstate disease burden relative to a standard population. Payer disparities are large (T2DM 2.96x, hypertension 2.58x higher in Medicare) but Medicare enrollment itself correlates with age and disability, so these ratios are confounded by age/eligibility structure, not necessarily payer effects per se.
What the model learned
The model achieves strong held-out discrimination (AUC 0.9220) and low Brier (0.01025), with only mild train-to-test drop (0.9358 to 0.9220) suggesting limited overfitting overall. The two dominant predictors by z-score are log(encounters) (OR 10.52) and age per decade (OR 10.96, with a negative quadratic term OR 0.899 producing the expected decelerating age effect) — encounters is the concern here: healthcare utilization volume plausibly increases *because* someone is dying (more visits, hospitalizations, escalating care near end of life), which is a classic label-leakage/reverse-causation risk rather than a clean upstream predictor. Active conditions count carries a counterintuitive protective OR (0.925, z=-3.07), likely reflecting collinearity with age/encounters rather than a true protective effect. Several clinically plausible predictors (heart failure OR 1.44, COPD OR 0.66, kidney disease OR 1.28, male sex OR 1.23) all have 95% CIs crossing 1, meaning we cannot reject no-association for any of them individually — they should not be presented as confirmed risk/protective factors.
Where this model would fail
Discrimination collapses in exactly the subgroup that matters most operationally: age 65+ AUC drops to 0.6693 (vs 0.9901 in 45-64, though that band has only 8 events and is likely unstable), and Medicare payer AUC is 0.6207 (51 events) versus 0.9982 for Aetna (8 events) — small-event subgroups generally should be read with caution regardless of the point AUC. At the chosen operating point (top decile flagged), the model catches only 55 of 79 true events (FN=24) while generating 612 false positives against 5971 true negatives — a precision of about 8%, meaning most flagged patients will not die. Calibration looks good in aggregate (decile 10: predicted 10.4% vs observed 8.4%) but decile 9 already shows overprediction is reversing to slight underprediction (2.25% -> 2.70%), and the near-zero deciles (1 and 6) are trivially well-calibrated only because events are essentially absent there — this tells us little about extreme-risk calibration stability.
What I would check next
First, audit whether "encounters (log)" is measured in a window that could include care driven by terminal decline — if so, this is leakage and the model's real prospective AUC would be lower. Second, rerun subgroup AUCs with event counts reported alongside every estimate (several subgroup AUCs above rest on single-digit event counts and are not reliable). Third, investigate the counterintuitive active-conditions OR for collinearity with age and encounters using variance inflation checks. Fourth, before any deployment, re-anchor the operating point to a clinically tolerable false-positive/precision tradeoff rather than a fixed top-decile cutoff, and treat all reported ORs and disparity ratios as adjusted associations only — none of this is causal evidence for intervention.

Synthetic corpus throughout — no real patients. The model is a demonstration of method and reporting discipline, not a clinical instrument: it has had no prospective validation, no calibration to a real population, and no fairness audit beyond the subgroup discrimination shown. Computed 2026-08-29 23:40:37; cached one hour.