φ(ai) PHI AI
DocumentationDocs
Repository

Model analytics & monitoring

Every model on the platform, watched: live performance metrics computed from the Data Store and the audit trail (never self-reported), drift recognition against the population each model learned from, and alerting when a threshold breaks. An unwatched model is an unmanaged one — this page is the management surface's other half, beside the Control panel's registry.

Open alerts

SeverityModel slotMetricAlert Fired
WARNING claims separation Score separation between denied and paid claims is 0 points on the validation sample - the model is barely discriminating. Review features. 2026-08-29 00:21
WARNING segmentation review_backlog 50 low-confidence classifications awaiting human review - withheld-pending-review content is accumulating. 2026-08-29 00:21
WARNING coding undocumented_billing 8 billed claims have no supporting note in their window - a compliance backlog, not just a finding count. 2026-08-29 00:21

Acknowledging is audited under your name; a condition that recurs after acknowledgment re-fires. Alerts never auto-close.

The registry, watched — every model

Complete coverage by construction: this table enumerates the model registry itself, so a model cannot exist without a monitoring row — including anything registered this session. Idle is a state, not a gap: a model with no invocations is watched and shown as idle.

ModelSlotStatusInvocationsLast activityOpen alerts
PHI RAG ACTIVE assistant enabled 0 idle 0 Configure
Claude Sonnet 5 ACTIVE assistant enabled 0 idle 0 Configure
Scripted fallback ACTIVE assistant enabled 0 idle 0 Configure
No-show risk model ACTIVE noshow enabled 0 idle 0 Configure
Sensitivity classifier ACTIVE segmentation enabled 0 idle 1 Configure
Scheduling optimizer ACTIVE scheduling enabled 0 idle 0 Configure
Terminology mapper ACTIVE ingest enabled 0 idle 0 Configure
Denial risk model ACTIVE claims enabled 0 idle 1 Configure
Prior auth evidence assembler ACTIVE priorauth enabled 0 idle 0 Configure
Coding integrity model ACTIVE coding enabled 0 idle 1 Configure
ROI requirements validator ACTIVE roi enabled 0 idle 0 Configure
Ambient listener — faster-whisper ACTIVE ambient enabled 0 idle 0 Configure
Inbox triage ranker ACTIVE triage enabled 0 idle 0 Configure
ROI requirements validator ACTIVE roi enabled 0 idle 0 Configure
Ambient listener — faster-whisper ACTIVE ambient enabled 0 idle 0 Configure
Inbox triage ranker ACTIVE triage enabled 0 idle 0 Configure
Imaging read assistant ACTIVE imaging enabled 0 idle 0 Configure
Patient instructions writer ACTIVE instructions enabled 0 idle 0 Configure
Ingest QA reviewer ACTIVE ingestqa enabled 0 idle 0 Configure

Assistant-slot models share the slot's question volume (ai.query names the model serving each one); every other slot counts its own model.inference events. Deep metrics for the slots with performance surfaces follow below.

Denial risk model — performance & calibration

8,740
pending claims scored
8.9 vs 8.9
mean score, denied vs paid (validation n=400)
0 pts
separation — the model's discrimination, measured against history
Score bandClaims Actually denied
med 320 9%
low 80 8%

Calibration read: a healthy model's bands order cleanly — the high band's actual denial rate should sit well above the low band's.

Drift recognition — input population vs baseline

0.001
payer-mix PSI (≥0.10 watch · ≥0.25 act) · window split 2026-02-07
9.1% → 9%
denial rate, baseline → current window
PayerBaseline share Current share
Aetna 13.8% 13.2%
Blue Cross Blue Shield 17.1% 17.1%
Cigna 12.2% 12.3%
Humana 7.7% 8.1%
Medicaid 17.1% 17.7%
Medicare 15.9% 15.1%
UnitedHealthcare 16.1% 16.4%

The other cores

ModelKey metrics, live
Prior auth evidence assembler 0 appeal drafts · 0 signed by a clinician (acceptance is the metric that matters — an assembler whose drafts nobody signs is noise).
Coding integrity model 15 under-coded findings open · 8 under-documented claims · 0 clinician queries drafted.
ROI requirements validator 0 validations run · 0 releases blocked on unmet requirements (each with the requirement and citation recorded).
Sensitivity classifier 50 low-confidence classifications awaiting review · 0 human reclassifications (audited).
Ambient listener (speech-to-text) 0 transcription sessions · 0 visit notes drafted · 0 visit-level consent attestations (capture never runs without the gate's verdict on the trail).
PHI RAG / foundation model 0 questions · 0 audited record reads across them.
What monitoring cannot do. These metrics watch the models; they do not validate them clinically, and they do not discharge the organization's AI-oversight duties — local validation before clinical use, subgroup review on the fairness screen, and the authority to pull the Control panel's switches remain human responsibilities.