Design Theory

When False Calculations Become Facts

Robodebt distributed aggregate income across periods that had not been observed. The process then treated the resulting figures as facts.

Macro view of a thickly painted oil surface: a massive dark brown block stands on a pale floor, with a single row of small identical slabs running through it. Behind the block they carry the same dark finish as the block itself; in front they are pale and hollow. A small indigo block stands exactly where the pale slabs turn dark.
Macro view of a thickly painted oil surface: a massive dark brown block stands on a pale floor, with a single row of small identical slabs running through it. Behind the block they carry the same dark finish as the block itself; in front they are pale and hollow. A small indigo block stands exactly where the pale slabs turn dark.

Australia’s Robodebt system derived income for individual fortnights from longer-term tax records by spreading the reported total evenly across them. It thus produced an exact debt from an income pattern the data did not contain. The Robodebt Royal Commission’s final report documents how that assumption became the basis for debts at scale.

The decisive mechanism was not the division alone. A discrepancy that had once prompted an administrative inquiry could be filled with substitute data and treated as a quantified debt in the online process. The Commonwealth Ombudsman’s 2017 report also marks the boundary: when actual income for every fortnight was entered, the same software could calculate correctly. The problem was what counted as sufficient evidence and what claim the result was allowed to support.

That transition is central to organizational uses of AI. A score, classification, or recommendation does not become a fact merely because a system returns an unambiguous value. It acquires institutional authority when a process allows that value to govern an investigation, benefit, debt, or other decision.

The necessary evidentiary boundary

“False” does not necessarily mean that the division was faulty. The tax records established an aggregate; Robodebt supplied its distribution over time and concealed which part of the result had been observed and which had been modeled. At the same time, an investigative lead became a debt: missing entries could be replaced by the average while the practical work of rebuttal moved to the affected person.

The evidence supports a tightly bounded claim. In an internal comparison of 229 cases, averaging almost never reproduced the manual calculation; the group was not identified as a representative sample, however, and deviations could also favor the individual. In Prygodicz v Commonwealth of Australia (No 2), apportionment without evidence of constant income could not establish a debt. This is not a general prohibition on averaging but a limit on claims whose evidence must be more temporally precise than the available data.

How a model output becomes a fact

Pseudo-precision here is a mismatch between the reach of the data and the reach of the claim. Several organizational steps turned the result into an enforceable debt: the tax authority supplied the total, the software distributed it, the process treated the result as an established liability, and the communication placed the burden of disproof on the individual. Reviewing the formula alone is therefore not enough. Its output’s status in the process must be reviewed too.

The same test applies when AI systems supply organizations with scores, rankings, or recommendations. At least four things must remain visible: the data actually observed, the assumptions added by the model, the uncertainty of the result, and the decisions for which it may be used. Without that separation, a plausible compression can quietly become an institutional fact.

Applying the same test to our own models

Intelligence Architecture also names layers, numbers steps, and uses playbooks. These forms can preserve distinctions and support collective work. They can also make a provisional judgment look measured.

The site’s foundation article distinguishes four activities: sense, interpret, decide, act. They are analytical tools, not a claim about a universal sequence. Actual work can loop, skip steps, or involve several activities at once. The four terms earn their place only when they change a concrete inquiry or make a break in the process visible. Their number proves nothing.

Readiness and maturity scores face the same question. Before an aggregate is allowed to describe an organization or direct investment, its weighting, missing data, borderline cases, and permitted uses must be visible. The critical review therefore starts before the arithmetic: what do the observations already support, what did the model add, and who turns the result into a reason for action?

Oliver Wrede writes and teaches on interface design, knowledge systems, and the architecture of intelligence in organizations. He is interested in how humans, institutions, and machines reason together — and how design shapes the quality of that reasoning.

More from Oliver Wrede