Epic sepsis model accuracy: 0.63 measured, up to 0.83 claimed. The number travelled and the caveat didn't
Jamie Watters
Operational resilience and AI delivery practitioner. Technology since 1985.

In June 2021 a team at Michigan Medicine published what happened when they ran Epic's sepsis prediction model against their own patients. Across 27,697 patients and 38,455 hospitalisations, the area under the curve was 0.63, with a confidence interval of 0.62 to 0.64. Epic's own figure for the same model was 0.76 to 0.83. The paper says where that figure came from: "internal documentation (shared with permission)", plus a conference proceeding co-authored with Epic that gave 0.73 (Wong et al., 2021).
Read those two numbers again, because the gap between them is the whole piece. One was measured by people with no stake in the answer, on a defined population, and published with its interval. The other lived in a document a hospital could see with the vendor's permission. By the time the model was running at what the same paper calls "hundreds of hospitals", only one of those numbers had travelled (Wong et al., 2021).
I don't have a better name for it, so I'm calling it hedge asymmetry. A number reaches you at full strength. The caveat attached to it reaches you weakened, or late, or as a parenthesis, or not at all.

The vendor's number, as it left and as it arrived. Nothing in between is dishonest. The caveat just doesn't survive the journey.
The claim and the caveat travel at different speeds
What the model actually did, from the same paper (Wong et al., 2021): at the recommended threshold it caught 33% of sepsis cases and missed 67%. It raised an alert on 18% of all hospitalisations. And of the 1,709 sepsis patients it missed, 1,030, or 60%, still got antibiotics in time, because clinicians had found them without it.
You couldn't see any of that from 0.76 to 0.83. All of it was decided by 0.76 to 0.83, because that's the number that got the model bought, switched on and left at the default threshold. The authors put the structural problem in a line: proprietary models arrive with "confidential, non-peer-reviewed model performance documents that may not accurately reflect real-world model performance" (Wong et al., 2021).
The model kept running. A separate study looked at two safety-net hospitals in Harris County, Texas, across 145,885 encounters covering the whole of calendar year 2023. Version 1 was still in production there, with a sensitivity of 14.7% at a six-hour window (Ostermayer et al., 2024). That is two and a half years after the Michigan result appeared in JAMA Internal Medicine. The number had been corrected in June 2021. The correction hadn't reached the emergency departments serving the patients least able to go anywhere else.

Published June 2021. Still running through 2023. Checked again in February 2026, against a benchmark from a meeting.
The correction that repeats the mistake
Epic revised the model. Reporting from corporate documents in October 2022 describes three changes (Ross, STAT, 2022). Hospitals are now told to train it on their own data. The definition of sepsis onset moved to a more standard one. And the model leans less on antibiotic orders as a predictor. That last one matters. Antibiotic orders as an input mean the model was partly detecting sepsis that clinicians had already diagnosed and started treating.
Then, in February 2026, the same lead author published a prospective validation of the revised model in JAMA Network Open. Four US health systems, 227,091 inpatient encounters, August 2023 to March 2025. Encounter-level AUROC ran from 0.82 to 0.92 depending on the site. No author is affiliated with Epic and no Epic funding is declared (Wong et al., 2026). The authors keep their own hedges in the conclusion: improved discrimination, but "high institutional variability, low positive predictive value, and high alert burden".
This is the paper you'd send someone as the fix. Independent authors, large sample, prospective, peer reviewed. So here is where the benchmark it compares itself against comes from. The introduction says Epic has reported AUROC values for the revised model "between 0.83 and 0.86 across 3 internal validation sites" (Wong et al., 2026). The citation for that sentence reads: "Tyler Sundberg and Elliot First, Epic Systems, Microsoft Teams meeting, February 25, 2025".
A Teams meeting. In the paper that exists because the last vendor figure was an internal document.
There's a second one. The methods say that "per developer recommendation", prediction scores in the first hour for patients without a blood count result were left-censored, which means excluded from the analysis (Wong et al., 2026).
Nothing here is improper. The authors disclosed both things exactly as they should, and disclosure is how I know about them. That's the point. The disclosure worked and the asymmetry survived it. A reader who goes to the primary source, which few readers do, still receives the vendor's number at full strength in the text and its provenance in a bracket. Five years after the 2021 paper warned about confidential, non-peer-reviewed performance documents, the vendor's performance claim inside the validating paper is a meeting.
Why the buyer's side should care
I've sat on that side. At HSBC I assessed vendors' continuity, disaster recovery and resilience claims as a condition of onboarding, and answered the same questions from our own clients in the other direction. The Epic record has the shape I recognise from that chair. The number arrives in a deck, a questionnaire answer or a configuration default. The qualifier arrives, if it arrives at all, in a footnote, a meeting, or a sentence beginning "in internal validation". The number is a better traveller than its caveat, and that is a property of the number, not of anyone's honesty.
And the downstream is not a footnote. The number decides whether the thing is bought, what threshold it runs at, and whether anyone asks "compared to what?" At Epic's threshold the model alerted on nearly a fifth of admissions. It missed two thirds of the cases it existed to find. And 60% of the patients it missed were being caught by clinicians anyway (Wong et al., 2021). Those three facts were available to anyone who ran the measurement. The buyer had the vendor's number instead.
The check
The fix is not "ask for evidence". Every assurance function already asks for evidence, and a Teams meeting citation is evidence, correctly disclosed. The fix is to make the caveat travel with the number, so that wherever the number appears, the reader sees what kind of number it is.
The provenance check, for any performance figure a supplier gives you
- Where is it published? A peer-reviewed paper, a regulator filing, a report you can cite by page. If the answer is "internal validation", "shared with permission", "a call with the vendor" or "a Teams meeting", the figure is a claim, not a measurement. Record it as one.
- Who measured it, and who paid? Independent of the vendor in authorship and in funding, or not. Financially independent is not the same as informationally independent: look for what the validators did "per developer recommendation".
- On what population, over what window, at what threshold? Retrospective at one site is not prospective at four. A figure without those three attached cannot be compared to anything, including the vendor's own earlier figure.
- Write the caveat into the number. Everywhere the number appears, in the deck, the risk assessment, the threshold configuration and the contract, it carries its answer to question 1. "0.83 to 0.86 (vendor, three internal sites, unpublished, cited to a meeting)" is a different input to a decision from "0.83 to 0.86" (Wong et al., 2026).
Pass rule: a performance figure enters a purchasing or go-live decision only in the form it takes after step 4. If step 1 cannot be completed, the figure does not enter the decision. It goes into the contract as something the vendor warrants, or it goes nowhere.
The objection
The objection from inside third-party risk is fair: we already caveat vendor numbers, and a citation to a Teams meeting is disclosure doing its job. The system worked.
Disclosure did work. That is exactly what I'm pointing at. The 2026 paper is as careful as a paper gets, and the vendor's unpublished number is still the benchmark its results are read against, in the sentence that says how much better the new model is. If the asymmetry gets through peer review at a JAMA journal, it gets through a procurement deck. The problem is not a missing caveat. It's that a caveat and a number don't move at the same speed, and every step between the measurement and the decision is a chance for them to separate.
Epic's revised model may well be better. Four health systems measured it and said so, with their own hedges attached (Wong et al., 2026). The benchmark that says how much better still comes from two Epic employees on a Teams call on 25 February 2025.
Sources
- Wong, A., Otles, E., Donnelly, J.P. et al. (2021) 'External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients', JAMA Internal Medicine, 181(8), pp. 1065-1070. The 27,697 patients, 38,455 hospitalisations, AUC 0.63 (0.62-0.64), Epic's 0.76-0.83 "in internal documentation (shared with permission)" and the conference proceeding's 0.73, sensitivity 33%, the 1,709 missed (67%) against 6,971 alerts (18%), the 1,030 (60%) treated in time anyway, "hundreds of hospitals", and the confidential-documents sentence. https://doi.org/10.1001/jamainternmed.2021.2626
- Wong, A., Currey, D., Schwinne, M. et al. (2026) 'Prospective validation of the Epic Sepsis Model version 2', JAMA Network Open, 9(2), e260181, published 27 February 2026. Four health systems, 227,091 encounters, 31 August 2023 to 11 March 2025, encounter-level AUROC 0.82 to 0.92, the conclusion's three hedges, the introduction's "between 0.83 and 0.86 across 3 internal validation sites" cited to "Tyler Sundberg and Elliot First, Epic Systems, Microsoft Teams meeting, February 25, 2025", and the methods' "per developer recommendation" left-censoring. Funding: the University of Michigan and NIH; no author affiliated with Epic and no Epic funding declared. https://doi.org/10.1001/jamanetworkopen.2026.0181
- Ostermayer, D. et al. (2024) JAMIA Open, 7(4), ooae133. 145,885 encounters across two Harris County, Texas safety-net hospitals during calendar year 2023; sensitivity 14.7%, specificity 95.3%, positive predictive value 7.6% at a six-hour window.
- Ross, C. (2022) STAT, 3 October 2022, reporting from Epic corporate documents on the revised model: local training, the sepsis-onset definition, reduced reliance on antibiotic orders.
- The provenance check is mine. It is not a script I have run for years; it is how I would put, in four lines, what a vendor assessment is trying to establish about a number. Standing: I assessed vendors' resilience claims at HSBC as a condition of onboarding and answered client due diligence in the other direction.
- Companion piece: four questions to ask any AI statistic before you repeat it, which sorted nineteen headline failure figures by what each one counts. This piece is about a different failure: a figure that counts something real and loses its caveat on the way to you.