Epic's sepsis model missed two thirds of cases. Its accuracy claim was never published.

Of the 1,709 sepsis patients Epic's prediction model failed to identify, 1,030 received timely antibiotics anyway.
Sixty per cent. The paper uses timely antibiotics as its proxy for clinicians having recognised the patient, which is not identical to a formal diagnosis but is what actually mattered at the bedside.
That single number does more damage to the deployment case than the accuracy figures do, and almost nobody quotes it.
What was measured
Wong and colleagues published an external validation in JAMA Internal Medicine in 2021. A retrospective cohort at Michigan Medicine: 27,697 patients, 38,455 hospitalisations, December 2018 to October 2019, with sepsis in 2,552 of them.
The model achieved an area under the curve of 0.63 at the hospitalisation level, confidence interval 0.62 to 0.64. AUC runs from 0.5, which is a coin flip, to 1.0, which is perfect.
Epic's claimed figure was 0.76 to 0.83.
Say the estimand out loud, because the two numbers may not be measuring the same thing. Wong's 0.63 is hospitalisation-level. Epic's internal figure almost certainly used a different endpoint and a different sepsis definition. That does not rescue the model, and the deployment-relevant numbers below are unaffected. But it sharpens the real complaint rather than softening it: the number hospitals bought against was not only unreviewed, it may not have been the number that describes use on a ward.
At the recommended threshold, sensitivity was 33%. Specificity 83%. Positive predictive value 12%. In the authors' own words, the model "did not identify 1709 patients with sepsis (67%) despite generating alerts for 6971 of all 38,455 hospitalized patients (18%)".
Miss two thirds of the cases. Alert on nearly one in five of everyone who comes through the door.

Where the vendor's number lived
This is the part that should bother you more than the gap itself.
The paper's own wording on provenance: the measured performance was "substantially worse than that reported by Epic Systems (AUC, 0.76-0.83) in internal documentation (shared with permission)".
Internal documentation. Shared with permission. The claim a hospital was buying against had never been through peer review, and could not be examined without the vendor's cooperation.
Wong and colleagues name the structural problem directly: "The increase and growth in deployment of proprietary models has led to an underbelly of confidential, non-peer-reviewed model performance documents."
That is not a story about one model being worse than advertised. It is a story about a class of claim that cannot be checked by the people carrying the risk.
The mistake, and it is not subtle
Among the model's predictors was whether antibiotics had been ordered.
Wong and colleagues did not publish the feature list, because they could not. That detail comes from Casey Ross's reporting at STAT on Epic's corporate documents, and from later academic work on the same mechanism.
Think about what that means. A clinician suspects sepsis, orders antibiotics, and the model reads that order as evidence of sepsis. It is partly detecting a diagnosis that has already been made by a human being.
That is label leakage. It is in every textbook. It inflates retrospective performance by construction, because in the training data the outcome and the predictor arrive together, and it deflates prospective usefulness for exactly the same reason.
Once that feature is in, the model can look better in validation while adding less on the ward.
Why nobody's list had it on it
I nearly wrote that the hospitals could not have caught this. That is wrong, and Epic said so at the time.
Responding to the paper, Epic stated that "the full mathematical formula and model inputs are available to administrators on their systems". Epic also defended the threshold, saying it "was relatively low and would be appropriate for a rapid response team that wants to cast a wide net".
And the hospitals demonstrably could catch it. Michigan Medicine did exactly that, which is why this paper exists. Harris Health did it later. You do not need the model's weights to measure local sensitivity, positive predictive value, calibration, alert burden or lead time. All of that is computable from your own data and your own outcomes.
So the honest version of the argument is narrower and, I think, worse.
The model arrived as a pre-trained component inside the electronic health record, shipped to hundreds of sites. Not as a system to be validated, but as a feature that was already there. The vendor had validated centrally and said so. Nothing about the deployment created a moment at which proving it worked in this population became somebody's named responsibility.
The architecture did not make validation impossible. It made validation look like it had already happened.
I want to be honest that this is the boundary case for that argument. Label leakage is a modelling error by any definition, and on the face of it Epic is a case where the modeller did make the mistake. My reading is that the leakage was baked into a product shipped generically at scale, which is an architecture decision rather than a project one, and the hospitals carrying the consequence had no route to it. That reading is arguable. If you reject it, the honest position is that the pattern holds in most of these cases rather than all of them.
Epic's redesign addresses the same two problems
In 2022, reporting from corporate documents identified three changes to the model.
Epic began recommending that hospitals train it on their own institutional data. It moved to a more commonly accepted standard for defining sepsis onset. And it reduced reliance on antibiotic orders as a predictor.
That is the generic-shipping decision and the leakage decision, both reversed by the vendor.
A redesign that addresses a criticism does not prove the criticism explained the whole performance gap. But it is hard to miss what changed.
One number to handle carefully
If you see the alert burden quoted, check which figure you are being given.
The paper reports a number needed to evaluate of 8 at the hospitalisation level, then 42, 59, 73 and 109 at 24, 12, 8 and 4 hour horizons. These are not competing estimates. They are the endpoints of a series, and the footnote says why: at the hospitalisation level each patient is evaluated only the first time the score crosses the threshold, while at each time horizon they are evaluated every time it does.
Anyone citing 109 has to say it is the repeated-alert, four-hour figure. Anyone citing 8 has to say it is the alert-once strategy. Quoting either without the qualifier is how a real finding turns into a talking point.
The bottleneck was never detection
The negative validation was published in June 2021. It was widely covered.
Version 1 of the model was still running through the whole of 2023 at two safety-net hospitals in Harris County, Texas, where a separate validation across 145,885 encounters found sensitivity of 14.7% and positive predictive value of 7.6% at a six-hour window, and found clinicians treating sepsis independently of the alert in roughly half of cases.
Two and a half years after publication. In safety-net emergency departments, among the patients with the least ability to go elsewhere.
Somebody detected this. Somebody measured it, wrote it up, and published it in a major journal. And the model kept running.
That gap between knowing and acting is not a technical problem, and no better model closes it.
Five years later, the same problem in the paper that validates the fix
Epic's version 2 has now been independently validated. JAMA Network Open, February 2026, four US health systems, 227,091 inpatient encounters. The revised model performs better, and I am not going to pretend otherwise.
The validation is financially independent of Epic. No author is affiliated with the company; the funding is university and National Institutes of Health.
It is not informationally independent, and the way it is not is the reason this article exists.
The methods state that "per developer recommendation, ESM v2 prediction scores within the first hour of evaluation for patients without an available complete blood count result were left-censored". The vendor shaped the analysis.
And then there is the benchmark the improvement is measured against. Epic's internal validation figures for version 2, cited in the introduction, are attributed like this:
(Tyler Sundberg and Elliot First, Epic Systems, Microsoft Teams meeting, February 25, 2025.)
Five years after Wong and colleagues named "an underbelly of confidential, non-peer-reviewed model performance documents", the vendor's performance claim, inside the independent paper validating the fix, is a Teams meeting.
So the model got better and the thing underneath it did not move at all.
Before you sign anything, the questions are not about the headline figure. Ask for the validation on patients like yours, at the threshold you will actually run, measured by somebody who does not work for the vendor. Ask what the alert burden is at that threshold, and who owns turning it off.
And ask where the number came from. If the answer is internal documentation shared with permission, or a meeting, you have not been given a measurement.