φ(ai) PHI AI
DocumentationDocs
Repository

Bulk import manager

FHIR Bulk Data Export from the source EMR — the only way to ingest an entire population rather than known patients. Kickoff, asynchronous status polling, NDJSON download, then encrypt-store-index for every record. The vendor's real seams are honored: on Epic a run needs a Group FHIR ID, is limited to once per 24 hours, and is always a full re-extract — there is no incremental mode.

116,339
records in the last clean run · Aug 27, 2026
eGrp-7ac2
group FHIR ID (from Source EMR configuration)
ready
a kickoff is permitted now
The run, its stages and its watermark behavior are recorded like the real pipeline's.

Run history

StartedFinishedSourceGroup StatusRecordsNotes
2026-08-28 01:00 → Epic eGrp-7ac2 failed 61,204 NDJSON download interrupted at file 14/22; watermark NOT advanced - a dirty run never moves the clean-run boundary.
2026-08-27 01:00 → 2026-08-27 02:20 Epic eGrp-7ac2 complete 116,339 Full population re-extract (Bulk Data Export has no incremental mode on this vendor); watermark advanced.
2026-08-26 01:00 → 2026-08-26 01:00 Epic eGrp-7ac2 refused 0 Second kickoff inside 24 hours - Epic rate-limits Bulk Data Export to once per 24h per group and client; refused rather than queued silently.
2026-08-25 01:00 → 2026-08-25 02:12 Epic eGrp-7ac2 complete 114,212 Full population re-extract; watermark advanced.
The watermark rule. A clean run advances the ingestion watermark; a failed or partial run never does — re-running re-extracts everything rather than trusting a boundary a dirty run cannot vouch for. The refused row above is the vendor's own 24-hour rate limit surfacing as a refusal instead of a silent queue.

Reconciliation — source EMR vs PHI AI vs target EMR

Three systems should agree about how many records exist — and where they deliberately do not, the difference must carry a named cause. An unexplained difference is the finding.

Counts by resource type across all three systems; recorded as integration.reconciliation.