Commentary

When Does a System Count as Intelligent?

Bernard Jennings proposes a test based on a system's internal architecture. Organizations still have to decide what authority, if any, its output should carry.

A close-up of a thickly painted oil painting: a precisely surveyed field outlined in corn yellow with dark blocks inside; from one corner, a thin green line crosses the open plane toward a small indigo block near the edge.
A close-up of a thickly painted oil painting: a precisely surveyed field outlined in corn yellow with dark blocks inside; from one corner, a thin green line crosses the open plane toward a small indigo block near the edge.

A basic classifier can wield considerable power when its threshold automatically removes a case from further review. A far more capable system could be confined to advice that nobody is obliged to follow. Organizational authority does not rise automatically with technical sophistication.

This distinction matters when we ask whether an AI system is intelligent. A strong definition may change how we describe the system. It cannot tell an institution who should act on the output, who may object, or who remains answerable for the result.

Bernard Jennings’s working paper The Architecture of Intelligence: Definition, Spectrum, and Ceiling makes an unusually demanding proposal about what should count as intelligence.1 Setting it beside the separate question of organizational authority shows both the value and the limit of an architectural definition.

A test aimed at internal organization

Many definitions begin with performance: Can the system solve problems, pursue goals, or succeed in different environments? Jennings looks instead at how the performance arises. In his account, an intelligent system must deal with demands on its own coherence. In plain language, it must register an internal conflict or gap as a problem for the system itself and take part in resolving it.

The resulting behavior must also contain what Jennings calls residue: something that cannot already be predicted from the system’s enduring capabilities together with the current input. Three conditions follow. The system must recognize the coherence problem as its own, carry out the work of resolving it, and produce residue in the process. Chance, novelty, and computational complexity do not qualify by themselves.

This is more than a switch from an output test to a diagram of components. Jennings links a claim about internal architecture to a proposed signature in behavior. When the architecture is only partly observable, applying the test becomes an empirical and interpretive task.

The evidence remains provisional

Jennings uses the definition to exclude chess engines, classical optimization systems, and language models that predate recent reasoning architectures. He argues that their outputs at inference time can be explained by capabilities already built into the system plus the current input.

His treatment of newer systems is more tentative. Reasoning models that run additional loops during inference, recursive architectures, and multimodal foundation models still receive a negative verdict, but Jennings describes their status as disputed because evidence about their internal operation is incomplete.

The publication status is important here. Zenodo identifies the paper as version 01 of a working paper. It proposes a theoretical criterion and sketches ways to falsify it; it does not demonstrate residue empirically in the contested systems. Jennings assigns that work to AI research, cognitive science, and the other fields he addresses.

His proposed ceiling depends on companion work. Eight Coherence Resolution Modes are meant to cover every way a system might resolve such a problem. If the list is incomplete, the ceiling derived from it cannot be fixed. The companion article on Jennings’s ceiling examines that dependency in detail.

Authority has to come from the organization

Even a positive result under Jennings’s test would settle only how the system is classified within his theory. The institution deploying it would still have to decide whether an output counts as information, advice, authorization, or an automatic trigger. Those choices determine its practical force.

The NIST AI Risk Management Framework gives this institutional work a concrete shape.2 The voluntary framework calls for clear risk-management roles and communication channels (GOVERN 2.1), executive responsibility for development and deployment risks (GOVERN 2.3), and distinct responsibilities for human–AI configurations and oversight (GOVERN 3.2).

For systems already in use, it also calls for ways in which users and affected communities can report problems and appeal outcomes (MEASURE 3.3). Other provisions address overriding or superseding a system’s output, removing the system from a process, and deactivating it (MANAGE 2.4 and 4.1).

NIST offers guidance rather than evidence that these arrangements work in every institution. Its relevance lies in the responsibilities it asks an organization to assign. People need the authority, information, and time to challenge an output. Someone must be able to change or stop its use. Technical design can support those powers, but the system cannot award them to itself.

An assessment should therefore keep the theoretical and organizational inquiries separate. For the system: What operations does it perform, what evidence supports calling it intelligent, and which conclusions remain open? For the deployment: Who sets the purpose, who may authorize action based on the output, how can an objection reach the decision in time, and who bears the consequences? Jennings offers a serious proposal for the first inquiry. Every organization still owes an answer to the second.

Footnotes

  1. Bernard Jennings, “The Architecture of Intelligence: Definition, Spectrum, and Ceiling”, Zenodo working paper, version 01, 22 May 2026. The record identifies the publication as a working paper. Jennings describes his definition as revisionary, draws the eight modes from the companion paper Coherence Resolution Modes (CRM) in Autonomous Systems, and leaves empirical tests of his criteria to future research.

  2. National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0), January 2023. The framework is voluntary, non-sector-specific, and use-case agnostic. The cited subcategories describe desired risk-management outcomes; they are not evidence that the proposed measures are effective in practice.

Oliver Wrede writes and teaches on interface design, knowledge systems, and the architecture of intelligence in organizations. He is interested in how humans, institutions, and machines reason together — and how design shapes the quality of that reasoning.

More from Oliver Wrede