Foundations

Stuart Russell and the Conditions for Useful Correction

The Off-Switch Game explains why an agent may benefit from waiting. Evidence from a working decision system shows what an organization must add.

A dark machine-like mass rests on a pale painted surface. Its connection to a small indigo block stops at a gap marked by a glowing burnt-orange contact.
A dark machine-like mass rests on a pale painted surface. Its connection to a small indigo block stops at a gap marked by a glowing burnt-orange contact.

The Off-Switch Game asks a narrow question: when does an optimizing system benefit from waiting for a human decision? In its baseline construction, waiting has value because it provides additional information about an uncertain outcome.1

That result does not give an AI system an organizational virtue. Human correction becomes effective only when the reviewer has relevant different information, may change the next step, and can transmit the objection through a reliable channel.

Why an optimizer might choose to wait

The original game, developed by Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Russell, contains one human H, one robot R, and a contemplated action a. R is designed to optimize the utility Uₐ of that action for H but knows its value only as a probability distribution. R can act, switch itself off for utility normalized to zero, or announce the action and wait. Only after waiting does H decide whether the action proceeds.

In the baseline, a rational H permits the action at non-negative utility and stops R at negative utility. Waiting is therefore at least as valuable as the best immediate option and becomes more valuable when R considers both positive and negative values possible. H’s response supplies information about the utility on which both parties are aligned; the game does not model a conflict between two objectives.1

The result depends on narrow conditions: R remains genuinely uncertain and H’s response is informative about that uncertainty. It establishes neither that a real reviewer knows the correct objective nor that an institution provides the necessary information and rights.

What the AFST case adds

A retrospective field study examined the Allegheny Family Screening Tool (AFST), which supported decisions about whether a child-welfare referral should be investigated. An out-of-home-placement subscore of 18 or more produced a mandatory-screen-in label; departing from it required a supervisor’s approval. A technical fault caused some real-time inputs to be calculated incorrectly.2

Researchers compared the visible score with one reconstructed from corrected data. A score appeared in 92.5 percent of the referrals analyzed. When displays were wrong, workers followed the visible number less closely and their choices aligned more closely with the reconstructed score’s recommendation. Only about two thirds of cases carrying the mandatory label were screened in. The retrospective study cannot experimentally isolate the reasons for individual decisions.

The organizational conditions are the important result. Screeners knew the referral call and administrative history, brought experience from the earlier workflow, and could override the output, sometimes with supervisory approval. Different evidence and decision rights together allowed them to offset some erroneous displays. Our companion article How a Risk Score Changed Hotline Work examines the AFST in more detail.

When the human knows less

The baseline gives H all the information R has about the choice. A later model, the Partially Observable Off-Switch Game, gives the two sides different information. It first appeared as a preprint in 2024 and was published at AAAI in 2025.3

In that setting, optimal play can lead R to avoid shutdown even though R pursues the common payoff and H is perfectly rational. More information or communication can improve expected joint utility or leave it unchanged, but cannot reduce it within the model. Even so, a limited channel can reduce the optimal level of deference. For a particular decision, the human may not be the better-informed party.

The later result narrows the first one rather than overturning it. Objective uncertainty does not guarantee corrigibility once knowledge is divided. The right to intervene matters, but so does the content that can pass between the parties.

When review can change the outcome

An effective review arrangement has to preserve useful differences between model and reviewer. The reviewer may know facts the model never received, interpret the same evidence from a different professional position, or see consequences that were outside the training target. The workflow must then give the reviewer enough time and a reliable way to change what happens next. Comment access is no substitute for an override.

The interface also determines what can be challenged. Relevant inputs, uncertainties, and operating limits need to be visible at a level that supports a specific objection. After the decision, records of overrides, errors, and outcomes must reach whoever can change thresholds and responsibilities. Otherwise review is an isolated checkpoint rather than a source of organizational learning.

Approval rates cannot establish whether this arrangement works. Agreement may reflect excellent recommendations, but it may also reflect copied information, automation bias, or weak authority. The operational question is which consequential errors intervention prevented, which ones it introduced, and what remained beyond the view of both parties.

Russell’s game contributes a precise piece of that account: uncertainty can make correction valuable to an optimizer. The Allegheny study shows the institutional conditions under which people were able to counter some erroneous outputs. The partially observable extension then shows how quickly the result changes when each side has only part of the evidence.

Footnotes

  1. Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell, “The Off-Switch Game”, IJCAI 2017; open preprint. The paper analyzes a deliberately small game and later extends the baseline to suboptimal human behavior. 2

  2. Maria De-Arteaga, Riccardo Fogliato, and Alexandra Chouldechova, “A Case for Humans-in-the-Loop: Decisions in the Presence of Erroneous Algorithmic Scores”, CHI 2020; preprint. Its evidence is a retrospective study of one operational setting, not a universal causal claim about human review.

  3. Andrew Garber, Rohan Subramani, Linus Luu, Mark Bedaywi, Stuart Russell, and Scott Emmons, “The Partially Observable Off-Switch Game”, 2024 preprint; subsequently AAAI 2025.

Oliver Wrede writes and teaches on interface design, knowledge systems, and the architecture of intelligence in organizations. He is interested in how humans, institutions, and machines reason together — and how design shapes the quality of that reasoning.

More from Oliver Wrede