Can You Actually Trust an AI?
The Horizon case shows why a technically consistent record cannot warrant its own truth. Trust requires a route from evidence to challenge and correction.

On 23 April 2021, the Court of Appeal in London quashed the convictions of 39 former Post Office branch operators and employees. They had been prosecuted for theft, fraud, or false accounting. Horizon accounting data had been essential to their prosecutions and convictions. The court found failures of investigation and disclosure and held that the prosecutions had been unfair and an affront to the public conscience.1
“Faulty software caused wrongful convictions” is an incomplete account of what went wrong. The same judgment dismissed three appeals because Horizon’s reliability was not essential to the evidence against those appellants. Electronic records had not become worthless, and the court did not find every Horizon entry false. The deeper failure was a procedure that allowed a disputed number to carry more authority than the people who lacked the information to challenge it.
The number answered a smaller question
Horizon recorded transactions and derived the cash and stock that should have been present in a branch. At the end of a trading period, the branch operator entered the physical totals. Horizon could then compute a difference exactly. That arithmetic answered whether the declared amount matched the balance derived from the system’s records.
It did not establish whether those records matched events in the branch, whether money was actually missing, why a difference had arisen, or who was responsible. Yet the operating process moved quickly beyond the smaller claim. A branch could not roll over into its next trading period without completing a statement based on Horizon figures. A shortfall had to be made good from personal funds or settled centrally as money owed to the Post Office.
There was no function inside Horizon for marking the figure as disputed. The branch operator had to use a telephone helpline, while the amount remained in the accounts. The courts later found that the Post Office treated challenged shortfalls as debts. In the cases ultimately overturned, a calculation could therefore become an allegation before its cause had been established.2
A trace has no authority of its own
The High Court’s 2019 Horizon Issues judgment exposed the asymmetry behind that conversion. For Legacy Horizon and the first online version, it found a material risk of errors arising during data entry, transfer, or processing. Horizon contained many controls, but automatic integrity mechanisms had failed on documented occasions; bugs and data errors had caused branch-account discrepancies.
Fujitsu personnel also held powerful roles capable of inserting, altering, or deleting transaction data and rebuilding branch data without the operator’s knowledge or consent. The court found the controls and records for privileged activity inadequate. Meanwhile, the Post Office possessed reporting facilities that branches did not. Some causes of discrepancies could not be investigated with the information available locally, leaving branch operators dependent on cooperation from the two organizations whose system and claims they were disputing.
Nor did the existence of detailed audit data settle the matter. The court did not suggest that full audit data had to be obtained for every transaction correction. It found, however, that access could be necessary when the cause or history of a disputed correction could not be resolved from other material, particularly in criminal or internal disciplinary proceedings. The value of the trace depended on access, use, and timing: who could retrieve it, whether the responsible decision-maker examined it, and whether an objection could still interrupt the process.2
The AI lesson is procedural
Horizon was not an AI system. Its history cannot establish how often a language model will fail. It does show why a technical output should not be allowed to certify the claim made on its behalf. With an AI-generated assessment, there are further transitions to examine. A source may support one sentence while the selection omits a decisive fact. Even an accurate summary may not justify the recommendation attached to it.
The procedure must therefore begin with the decision being prepared and the consequences of error. Material claims should lead to retrievable, versioned evidence. Independent reviewers need access beyond the interface and summary they are checking. Privileged activity must be controlled and recorded, while performance in use is observed. A challenge needs to reach an accountable person before the outcome becomes irreversible; authority to override, suspend, and communicate an incident has to be assigned in advance.
This is close to the work described by the voluntary NIST AI Risk Management Framework. NIST treats trustworthiness as a combination of characteristics managed across the lifecycle. The framework calls for independent review, routes through which users and affected communities can report problems and appeal outcomes, and mechanisms to override or deactivate systems and respond to incidents.3 It organizes responsibility; it does not certify a result.
Evidence may still be incomplete. Reviewers who are formally independent can share the same mistaken premise, and an appeal route can exist without the power to change anything. Confidence must remain specific to a use and open to withdrawal.
A good procedure guarantees neither truth nor fairness. It can make error discoverable, contestable, and correctable. That is enough to distinguish warranted trust from deference to a polished output.
Footnotes
-
Court of Appeal (Criminal Division), Hamilton & Others v Post Office Limited, [2021] EWCA Crim 577, especially paras. 1–5 and 447. The court allowed 39 appeals on both grounds and dismissed three. ↩
-
High Court of Justice, Bates & Others v Post Office Limited (No. 6: Horizon Issues), [2019] EWHC 3408 (QB), especially paras. 905–913 and 963–1030. The findings concern the versions and periods examined in the litigation; the court expressly cautioned against applying them wholesale to Horizon as it stood in December 2019. ↩ ↩2
-
Elham Tabassi, Artificial Intelligence Risk Management Framework (AI RMF 1.0), National Institute of Standards and Technology, 2023, especially MEASURE 1.3 and 3.3 and MANAGE 2.4, 4.1–4.3. ↩

