Artificial Intelligence

When a Score Becomes a Rule

Once an AI output is tied to thresholds, defaults, and authority, it changes the decision system.

A close-up of a thickly painted oil painting: a tool wall with neatly hung wrenches, one slot outlined in burnt orange; the small indigo block does not hang on the wall — it sits below, embedded in a network of recessed channels that visibly rearrange themselves around it.
A close-up of a thickly painted oil painting: a tool wall with neatly hung wrenches, one slot outlined in burnt orange; the small indigo block does not hang on the wall — it sits below, embedded in a network of recessed channels that visibly rearrange themselves around it.

At Allegheny County’s child welfare hotline, 18 became a policy boundary.

In August 2016, the county began using Version 1 of the Allegheny Family Screening Tool for one class of child-welfare referrals. The system produced two risk scores for each child. The operationally consequential one was a placement-risk score from 1 to 20. At 18, it entered the top 15 percent of the estimated risk distribution and triggered what the county called a mandatory screen-in. The label was stronger than the rule: a supervisor could still override it. But the route through the decision had changed.

I use AI here in the broad organizational sense: a system that derives a machine-made assessment from data. This is an older case. Version 1 used a logistic regression, not a large language model or an autonomous agent. The details below concern that version. In December 2018, the county introduced Version 2 with an updated algorithm, data sources, and policies. The older case separates the organizational issue from the novelty of current AI. A prediction acquires formal weight when it is connected to a threshold, a workflow and a set of permissions.

A score acquired a job in the organization

The AFST did not determine whether the allegation on the phone was true. It estimated the likelihood of an out-of-home placement within two years after an investigated referral. The model was trained only on referrals that had in fact been investigated. Its inputs came from linked administrative records concerning people associated with the referral. The allegation itself was not encoded. A call worker could therefore act on an acute incident described in the allegation that the score might miss.

This distinction caused a practical problem of language. Staff were accustomed to using “risk” for imminent harm. The score referred to a different future outcome. The county’s implementation evaluation shows why that boundary remained difficult to communicate during implementation. A model can be statistically specific while the word attached to its output remains organizationally ambiguous.

The number only became consequential through policy. For non-mandatory GPS referrals, the score was one input into the screener’s recommendation. At 18, the default flipped. A referral was expected to proceed to investigation unless a supervisor intervened. The machine did not issue a final decision; the procedure assigned extra force to its output.

The exception was used. Data from the first year showed supervisors overriding nearly one in four mandatory screen-ins. Screening decisions also corresponded more strongly with workers’ own risk-and-safety assessments than with the AFST score. The threshold altered the formal route without silencing professional judgment.

Defaults matter because they allocate effort and justification. The prescribed route required no override. Departure required someone with the appropriate authority to notice the case, disagree and authorize the exception. The override therefore does not cancel the structural effect of the threshold. It is part of that effect.

The intervention was larger than the model

The county also changed the division of work. Before implementation, hotline screeners mainly gathered information and passed it to supervisors, who made the screen-in or screen-out decision. Screeners could offer their view, but the decision belonged to the supervisory role. Under the new process, screeners completed risk and safety ratings and generated the family screening score. For non-mandatory GPS referrals, they made a recommendation that supervisors approved or changed.

The AFST was therefore introduced alongside a redistribution of judgment. The county overview describes the tool as additional information used with clinical judgment. That description is accurate, yet incomplete as an account of the intervention. The working system included the prediction target, the administrative data, the 1-to-20 display, the threshold, the exception rule, the screener’s new recommendation and the supervisor’s approval.

Treating the model as a self-contained tool hides these connections. It also encourages the wrong evaluation question. Asking whether screeners “used the tool” says little about whether the new arrangement produced better decisions. It does not reveal whose work increased, whether borderline cases received better attention, how often overrides were warranted or whether errors fell unevenly on different groups.

A useful case, with awkward evidence

Allegheny County should not be presented as a controlled demonstration of what predictive models achieve. Process rules and roles changed at the same time as the model was introduced. Any later difference in practice could reflect the score, the new responsibility given to screeners, supervisory behavior or their interaction. The case provides unusually concrete evidence about implementation. It does not isolate a model effect.

The setting also resists a tidy success narrative. A child protection investigation may prevent grave harm, and it may impose a serious burden on a family. Government agencies tend to hold more records about people who rely on public services. The researchers’ published case study therefore examines the possibility that administrative data could disadvantage communities already exposed to more state scrutiny. Human judgment is not a neutral benchmark here, but neither is historical data.

The same paper reports two errors in the original validation procedure. The same children, or siblings from the same referral, had crossed the training and test split; feature selection had also been performed before the split. Both made the reported performance estimates too optimistic. The researchers later corrected the validation procedure and recalculated the estimates. That did not alter the model already in use. Another difficulty remained: the model had been trained and evaluated on referrals that were screened in. Its intended use was to help choose among all referrals, including those for which later placement outcomes were not observed in the same way. This “selective labels” problem limited what the validation could establish.

None of these points proves that the AFST should never have been used. They show why model accuracy cannot carry the full argument for deployment. The claim being predicted, the population on which it was tested and the policy attached to the result all need separate scrutiny.

Design the decision system before selecting its component

A pilot involving a system that classifies, ranks, recommends or triggers action should start with a description of the decision arrangement it will alter. What exactly does the output claim? Which information is absent? Who sees it, and at what point? Does a score merely inform a judgment, change a default or trigger an action? Who may override it, and what must that person record? Which role gains responsibility, and which role may lose the chance to form an independent view?

Evaluation should follow the same boundary. Model performance belongs in it, but so do overrides, missed cases, distributional effects, changes in workload and the quality of the final decisions. If policies and roles change during the pilot, their contribution must be documented rather than credited to the model by default.

Eighteen did not decide anything by itself. It mattered because the procedure changed direction at that point. The useful question for an AI pilot is therefore not whether the model is merely a tool. It is what rights its output acquires in the workflow.

Oliver Wrede writes and teaches on interface design, knowledge systems, and the architecture of intelligence in organizations. He is interested in how humans, institutions, and machines reason together — and how design shapes the quality of that reasoning.

More from Oliver Wrede