What Must Be Decided Before Choosing an AI Model
Three bodies of theory with very different standing clarify the work that comes first: setting purpose, detecting material deviation, and making correction effective.

A system may include a shutdown function while leaving everyone around it unable to use that function well. The relevant evidence might arrive too late, the reviewer might lack authority, or the technical intervention might fail to alter an action already under way. Corrigibility is therefore more than a model property; it depends on the system in which the model is deployed.
Three bodies of work help inspect that deployment, though they carry very different weight. Stuart Russell formalizes an incentive to remain open to correction. Stafford Beer offers the organizational account needed to locate that correction. Bernard Jennings contributes a recent and speculative challenge to the language of ever-increasing intelligence. Bringing them together is the argument of this essay; the authors did not develop a common theory.
What an off-switch can and cannot establish
Russell’s criticism of the AI “standard model” concerns agents that optimize fixed objectives. If a formal target is a poor proxy for what people intend, increasing capability can magnify the mismatch by pursuing that proxy more effectively. The claim is conditional. Capability is not the problem in itself; a fixed objective that cannot be revised in light of observed effects is.
The 2017 Off-Switch Game gives a robot uncertainty about human utility and lets human behavior provide evidence about it. Under the model’s assumptions, the robot can have an incentive to preserve the option of shutdown.1 This formal result does not show that a deployed agent is corrigible, or that a person present in the workflow has meaningful control.
A partially observable version published in 2025 gives the parties unequal information. In some configurations, optimal agents avoid shutdown even though the human acts with perfect rationality.2 The difference matters operationally. A right to stop needs timely evidence, an effective means of intervention, and a role authorized to use it.
Beer’s account of where correction lives
The Viable System Model describes functions required for an organization to remain viable. Operational units respond to local conditions. One function coordinates their interdependence; another monitors current operations, allocates resources, and retains a direct audit channel. A separate function attends to the environment and the future. Policy reconciles present performance with adaptation and the organization’s identity.3
Beer treats these as recurring functions rather than departments. An operational unit can be analyzed as a viable system at its own level. Local autonomy is thus combined with wider coordination. A central body cannot process every detail, but filtering creates its own danger: a report that removes the anomaly may remove the reason to intervene. Direct checks and channels that carry evidence of consequential deviations are part of the regulatory design.
Beer did not apply the VSM to contemporary AI. In such a deployment, a model output becomes consequential through data access, interface choices, workflow rules, and delegated authority. An agent that initiates an action still operates with tools and limits assigned by the organization, which remains responsible for that allocation.
This does not imply manual approval for every output. Scope limits can reduce the failures that oversight must detect. Domain expertise, samples, and automated monitoring each reveal different kinds of deviation. Their adequacy can only be judged against the setting and the cost of being wrong.
Russell’s problem now has a place in the larger arrangement. Leaders responsible for policy set the purpose, but evidence from operations must be able to prompt them to revise it. A benchmark measures success on a specified task. It cannot decide whether the task is still a good proxy or whether its errors remain acceptable in use.
What Jennings adds, provisionally
Jennings published his independent manuscript in May 2026. It defines intelligence architecturally: a system must register and resolve coherence demands as its own, and its behavior must contain residue beyond what stable competence plus input would predict. From those premises, the paper proposes a spectrum with a floor and a ceiling.4
The modest contribution here is a distinction between greater operational capability and a different cognitive architecture. Faster processing and greater throughput may substantially improve performance without demonstrating the architectural change Jennings describes. Compute can expand the set of tasks solved in practice even if his classification remains unchanged.
The proposed ceiling is conditional on a separate completeness argument. The paper is recent, independent, and not an empirical confirmation of an upper bound on intelligence. Its vocabulary can sharpen a question about where improvement comes from. It is not mature enough to determine a governance decision.
The decisions that precede model choice
This ordering follows an organizational concern and does not rank the authors’ contributions as theories of cognition. For a deployment, the organization first needs a purpose and a named authority able to revise it when a proxy begins to displace the intended outcome. Evidence from use must reveal material departures soon enough to change course. Those revision rights must be backed by technical controls that can narrow, reverse, or stop the system while intervention can still matter.
Model selection belongs inside that prior design. A capable model cannot supply the purpose, the means to surface contradictory evidence, or the responsibility for acting on what that evidence reveals.
Footnotes
-
Stuart Russell, “Human-Compatible Artificial Intelligence”, in Human-Like Machine Intelligence (Oxford University Press, 2021), and Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell, “The Off-Switch Game”, Proceedings of IJCAI-17, 2017, 220–227. ↩
-
Andrew Garber, Rohan Subramani, Linus Luu, Mark Bedaywi, Stuart Russell, and Scott Emmons, “The Partially Observable Off-Switch Game”, Proceedings of the AAAI Conference on Artificial Intelligence 39(26), 2025, 27304–27311. ↩
-
Stafford Beer, Brain of the Firm (1972), The Heart of Enterprise (1979), and Diagnosing the System for Organizations (1985); see also his primary article “The Viable System Model: Its Provenance, Development, Methodology and Pathology”, Journal of the Operational Research Society 35 (1984), 7–25. Applying the model to contemporary AI systems is this essay’s synthesis, not Beer’s claim. ↩
-
Bernard Jennings, “The Architecture of Intelligence: Definition, Spectrum, and Ceiling”, independent working paper, version 01, published 22 May 2026. The abstract explicitly makes the proposed ceiling conditional on the completeness argument. ↩

