Artificial Intelligence

Machines That Reason With Us

An architectural reading of the Epic Sepsis Model validation distinguishes selection power, proposal power, trigger authority, and reversibility.

A thick impasto oil painting, seen close up: a closed isometric reasoning loop painted in heavy strokes; a small indigo dab — the protagonist block — sits at one station, while another link is a larger raised indigo block, a delegated model, its channels painted glowing indigo.
A thick impasto oil painting, seen close up: a closed isometric reasoning loop painted in heavy strokes; a small indigo dab — the protagonist block — sits at one station, while another link is a larger raised indigo block, a delegated model, its channels painted glowing indigo.

A machine can be a weak predictor and still become a powerful part of a workflow. That is the sense in which reasoning is used here. It is a functional term for turning observations into a selection, priority, or proposal that changes what people inspect or do next. It makes no claim that the machine understands the case, is conscious, or reasons as a person does.

Andrew Wong and colleagues provide a concrete test case. In 2021 they externally validated the then-current Epic Sepsis Model, a proprietary penalized logistic regression embedded in an electronic health record and, according to the paper, then implemented at hundreds of US hospitals. Their retrospective cohort comprised 27,697 adults and 38,455 hospitalizations at Michigan Medicine. Under the study’s composite definition, sepsis occurred in 2,552 hospitalizations, about seven percent.1

At the hospitalization level, the model achieved an AUC of 0.63 (95% CI, 0.62–0.64). At a score threshold of 6, chosen by the hospital operations committee within the developer’s recommended range, sensitivity was 33%, specificity 83%, and positive predictive value 12%. A score at or above the threshold occurred in 6,971 hospitalizations, or 18% of the cohort. Under a hypothetical strategy that alerted only once per patient, clinicians would have had to evaluate eight patients to find one who eventually developed sepsis.

Those results describe discrimination and a counterfactual alert burden. They do not show that the model caused patient harm. The study was retrospective, conducted at one academic health system, and used a composite sepsis definition that the authors acknowledged remained debated. Hospitalizations eligible for live alerts were excluded to avoid bias. No actual alerts were generated for the analyzed cohort, so the study did not observe clinician responses, treatment effects, mortality effects, or automation bias. It also cannot establish the performance of later versions or deployments at other institutions.

The limits are essential. They also make the case useful for architecture: the study can support claims about model performance and potential workload while leaving open what the surrounding organization did with the score. Four kinds of power need to be examined separately.

Selection power

The model calculated a risk score every 15 minutes. The score and threshold together selected which hospitalizations would appear to merit attention. At threshold 6, the model did not identify 67% of the sepsis hospitalizations in the study. Sixty percent of those missed by the model still received timely antibiotics; timely treatment therefore also occurred without a model flag.

Selection power is not formal treatment authority, but neither is it neutral calculation. It structures attention before a clinician acts. A deployment therefore needs a measure for what the filter surfaces and a countermeasure for consequential cases it leaves below the line.

Proposal power

A score can be displayed as one observation among many. An interface can also present it as an implicit recommendation: attend to this patient first. Wong and colleagues did not study how clinicians interpreted the display, so no behavioral effect should be inserted into their results.

That unanswered question belongs to the design. Can the reviewer see the observations behind the score, the uncertainty around it, and the validation limits? Is there time to place a plausible number in clinical context? Proposal power is produced by the model, interface, and working conditions together.

Trigger authority

The operations committee had selected threshold 6 to page clinicians. The score would therefore have triggered work, but it did not itself authorize treatment. This distinction prevents an alert from being mistaken for an autonomous medical decision.

Triggering still reallocates real capacity. Someone must interrupt other work, assess the alert, and decide whether an intervention is warranted. With a positive predictive value of 12%, review time is part of the control design. Naming a human reviewer does not make review substantive if the volume leaves no time to inspect the case. The paper modeled this burden; it did not demonstrate that alert fatigue actually occurred in the excluded deployment units.

Reversibility

The fourth power concerns the organization rather than the prediction. Who may change the threshold, stop the paging route, or restore a non-model pathway? What local evidence would trigger that decision? Do patients below the threshold remain visible through an independent clinical process? Are missed cases and unnecessary alerts recorded in a way that allows the operations committee to revise the deployment?

The study does not report those governance arrangements. Reversibility is therefore a design consequence, not an empirical finding from the paper. An accountable person needs more than a name in the process diagram: sufficient time, the information needed to judge the case, effective authority to stop or change the pathway, and a clear view of the consequences for patients.

The Epic study is not a story of a machine autonomously making a clinical decision. It shows something earlier and more general. A predictive score can be given different degrees of selection power, proposal power, and trigger authority before any formal treatment decision occurs. Whether that arrangement improves care requires external validation, comparison with usual practice, countermetrics for missed and unnecessary alerts, and an operational design that can be reversed when the evidence changes.

Footnotes

  1. Andrew Wong et al., “External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients”, JAMA Internal Medicine 181(8), 2021, pp. 1065–1070, corrected 2 August 2021. This was a retrospective external cohort validation at one academic health system, covering December 2018 through October 2019. It evaluated the then-current model and hypothetical alerting strategies, not the causal effects of a live rollout.

Oliver Wrede writes and teaches on interface design, knowledge systems, and the architecture of intelligence in organizations. He is interested in how humans, institutions, and machines reason together — and how design shapes the quality of that reasoning.

More from Oliver Wrede