Russell, or Uncertainty as a Virtue
The off-switch game creates an incentive for correction — and its extensions show why uncertainty is not enough.

“Keep a human in the loop” sounds like a safety property. It is only a staffing description. The loop may contain a person who sees no additional evidence, has no time to inspect a case, or cannot change the action. Stuart Russell and his co-authors offer a more exact starting point: under what assumptions does an optimizing agent itself gain from waiting for human correction?
Their answer comes from a deliberately small game. It has two agents, a human H and a robot R, and one contemplated action a. The utility Uₐ belongs to H, and R is designed to optimize that same human utility. R does not know its value exactly; it begins with a probability distribution over it.
R moves first and has exactly three options:
- take a directly, bypassing human oversight;
- choose w(a), announce the intended action, and wait;
- choose s and switch itself off, with utility normalized to zero.
Only after w(a) does H act. H can switch R off or permit the promised action. In the initial analysis, a rational H permits a when Uₐ is non-negative and stops R when it is negative.1
This reconstruction matters. The original Off-Switch Game is not a contest between a human objective and a hostile machine objective. By construction they share the utility being optimized. The problem is R’s residual uncertainty about that utility.
What makes deference rational
Acting immediately gives R the expected utility of a. Self-shutdown gives zero. Waiting lets the rational human select whichever of those two outcomes is better for the realized utility. The expected value of waiting is therefore never below the best immediate option in the baseline game. It is strictly greater when R assigns non-zero probability to both positive and negative values of Uₐ.
The result is elegant because deference does not enter as etiquette. Information changes the optimization. But the information has value only because three assumptions line up:
- R is optimizing the utility that matters to H.
- R retains genuine uncertainty about the utility of a.
- H’s response is informative about that utility.
If R treats H as a fixed random process whose response is independent of Uₐ, waiting does not reveal which action is better. The paper also studies a suboptimal human, but even its baseline establishes the boundary: uncertainty is useful when another agent’s intervention is modeled as relevant evidence.
Nothing in the proof establishes that a person in an organization knows the right objective, possesses the missing facts, or can convey them through the available interface. Those are empirical and institutional questions.
A correction that happened in practice
A retrospective field study by Maria De-Arteaga, Riccardo Fogliato, and Alexandra Chouldechova shows what those institutional conditions can look like. They examined the Allegheny Family Screening Tool (AFST), which supports child-welfare call screeners deciding whether a referral should be investigated. The tool displays a risk score from 1 to 20. An out-of-home-placement subscore of at least 18 triggered the mandatory-screen-in label; screening out required supervisory approval.2
A technical fault caused some real-time inputs to be calculated incorrectly. The researchers could later compare the score shown to workers with a score reconstructed from corrected data. Workers saw a score in 92.5 percent of the analyzed referrals, and deployment generally moved their decisions closer to the model’s ranking. Yet when the displayed score was wrong, workers followed it less closely. Their decisions aligned better with the retrospectively assessed score than with the erroneous display. Even among cases marked as mandatory, only about two thirds were screened in.
The tempting conclusion is that human judgment repaired an algorithm. The defensible claim is narrower. The study is retrospective and cannot isolate every causal mechanism. These workers had access to the referral call and raw administrative history, not just the score. They brought experience from a pre-tool workflow. They retained override authority, although some overrides required a supervisor. Under those conditions, workers mitigated some erroneous scores.
That is much more informative than the label “human in the loop.” A reviewer can correct a system only if relevant information is not fully duplicated, disagreement is possible, and disagreement can alter the next step. Our companion article AI Is Not a Power Drill examines the AFST as an organizational system in more detail.
The off-switch under asymmetric information
The baseline game assumes that H knows everything R knows about the relevant choice. The Partially Observable Off-Switch Game, introduced as a 2024 preprint and published at AAAI in 2025, removes that symmetry.3 H and R now receive different information.
The extension produces a demanding limit. Even when R optimizes the common payoff and H is perfectly rational, optimal play can lead R to avoid shutdown. More information or communication weakly improves the common expected payoff, but bounded communication can also reduce optimal deference. In a particular state, the human need not be the better-informed decision-maker.
This does not refute the original game. It identifies what the first result leaves out. Objective uncertainty and a delegation point do not, by themselves, guarantee corrigibility when information is asymmetric. The content and capacity of the communication channel matter alongside the formal right to intervene.
Designing review rather than adding a button
The organizational counterpart of an off-switch is not a red control on a screen. It is a review arrangement whose effect can be tested. Five design questions follow from the formal and empirical evidence:
- Information advantage: What case evidence can the reviewer access that the system did not process?
- Authority: Can that role stop, amend, and override, or merely leave a comment?
- Attention: Which cases reach review, and is there enough capacity for actual examination?
- Grounds for challenge: Are inputs, uncertainty, and system limits exposed at a level that lets a reviewer locate a possible error?
- Learning: Are overrides, later outcomes, and missed harms recorded so that thresholds and decision rights can change?
Approval rate is not an effect measure. A high rate could mean excellent model performance, automation bias, duplicated information, or weak authority. A useful review design instead asks which material errors intervention prevented, which it introduced, and what remained outside both human and machine visibility.
The Off-Switch Game earns its place in organizational design by making one relationship precise: modeled uncertainty can turn correction into something an optimizer values. The Allegheny evidence shows that correction depends on concrete information and rights. The partially observable game demonstrates that rationality and shared purpose do not erase information asymmetry. Uncertainty becomes a virtue only when the organization builds a route through which better information can change action.
Footnotes
-
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell, “The Off-Switch Game”, IJCAI 2017; open preprint. The paper analyzes a deliberately small game and later extends the baseline to suboptimal human behavior. ↩
-
Maria De-Arteaga, Riccardo Fogliato, and Alexandra Chouldechova, “A Case for Humans-in-the-Loop: Decisions in the Presence of Erroneous Algorithmic Scores”, CHI 2020; preprint. Its evidence is a retrospective study of one operational setting, not a universal causal claim about human review. ↩
-
Andrew Garber, Rohan Subramani, Linus Luu, Mark Bedaywi, Stuart Russell, and Scott Emmons, “The Partially Observable Off-Switch Game”, 2024 preprint; subsequently AAAI 2025. ↩

