Artificial Intelligence

When a Sepsis Score Becomes an Alert

A hospital committee turned a prediction into a page. External validation shows why every step from score to clinical work needs deliberate design.

A thick impasto oil painting, seen close up: a closed isometric reasoning loop painted in heavy strokes; a small indigo dab — the protagonist block — sits at one station, while another link is a larger raised indigo block, a delegated model, its channels painted glowing indigo.
A thick impasto oil painting, seen close up: a closed isometric reasoning loop painted in heavy strokes; a small indigo dab — the protagonist block — sits at one station, while another link is a larger raised indigo block, a delegated model, its channels painted glowing indigo.

A 2021 external validation of the then-current Epic Sepsis Model retrospectively calculated which Michigan Medicine hospitalizations the model would have flagged. The model produced a new score every fifteen minutes; an operations committee had selected 6, within the developer’s recommended range, as the point at which a clinician should receive a page.1

At that threshold, 6,971 of 38,455 hospitalizations, or 18 percent, would have been flagged. Only 12 percent of them later met the study’s sepsis definition, while 67 percent of later sepsis cases remained below the line. The retrospective study did not observe real pages or clinical responses.

The organizational issue is therefore visible before the clinical chronology: the model did not make a treatment decision but produced a possible review queue. Only the connection among threshold, page, available attention, and decision authority turned a score into an intervention in clinical work.

What the numbers say about the review queue

The then-current model was a proprietary penalized logistic regression embedded in the electronic health record. The study covered 27,697 adults and 38,455 hospitalizations; its composite definition identified sepsis in 2,552 of them. An AUC of 0.63 (95 percent confidence interval, 0.62–0.64) was a precise estimate of weak discrimination between hospitalizations with and without the later outcome.

The threshold made the work implications concrete. Of 100 hospitalizations with sepsis, 33 would have been flagged; of 100 without it, 83 would have stayed below the line. Roughly eight flagged hospitalizations would have required review for one eventual sepsis case. These measures cannot be collapsed into one “accuracy” figure: they describe both a demand on scarce clinical attention and a large group of unflagged cases.

Sixty percent of the sepsis hospitalizations missed by the model still received timely antibiotics, so many reached treatment without a model flag. That figure alone does not show whether or how sepsis was recognized, but it does show why an evaluation must include the clinical paths below the threshold.

How a score became a page

The Epic Sepsis Model recalculated a patient’s risk score every fifteen minutes. An operations committee chose 6, within the developer’s recommended range, as the threshold. At 6 or above, the system would page a clinician.

The threshold first selected which hospitalizations received additional attention. The interface then shaped how the clinician encountered the score: it could present the number as one observation among many or make it appear to be an instruction to examine this patient first. The validation study did not observe how clinicians read the display, so it provides no evidence that they deferred to it.

The page added a concrete action. It interrupted someone and requested a review, although the clinician still decided whether any intervention was warranted. Review consumes time even when it ends with no treatment. The operational effect therefore began before the treatment decision.

Governance determines whether that sequence can be changed. Someone needs the authority and monitoring data to move the cutoff, suspend alerts, or restore a non-model process when local evidence no longer supports the deployment. Merely naming an owner does not create that capacity.

What a live pathway still has to answer

Hospitalizations exposed to possible real ESM alerts were excluded from the analyzed cohort to reduce bias. The paper therefore cannot establish effects on treatment or mortality, and it provides no evidence of automation bias. Its authors also note the limits of one academic center and a composite sepsis definition that remained contested. Later model versions and deployments elsewhere require their own evidence.

The same boundary applies to alert fatigue. The calculations indicate a large potential workload; they do not show that staff in the excluded live units actually became fatigued. With a positive predictive value of 12 percent, the design must account for review capacity. Human oversight is meaningful only when clinicians have enough time and information to assess the case.

Wong and colleagues do not describe whether Michigan Medicine could suspend the alerts or how missed and unnecessary alerts informed governance. Evidence of benefit would need to compare the full clinical pathway with usual practice, measure outcomes on both sides of the threshold, and retain another way to identify patients whose scores remain below 6. Until such evidence exists, the retrospective score performance and the live clinical intervention remain different objects of evaluation.

Footnotes

  1. Andrew Wong et al., “External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients”, JAMA Internal Medicine 181(8), 2021, pp. 1065–1070, corrected 2 August 2021. This was a retrospective external cohort validation at one academic health system, covering December 2018 through October 2019. It evaluated the then-current model and hypothetical alerting strategies, not the causal effects of a live rollout.

Oliver Wrede writes and teaches on interface design, knowledge systems, and the architecture of intelligence in organizations. He is interested in how humans, institutions, and machines reason together — and how design shapes the quality of that reasoning.

More from Oliver Wrede