Artificial Intelligence

How a Risk Score Changed Hotline Work

Allegheny County tied the AFST to a default rule requiring investigation and reorganized screening around that rule. Policy made the model consequential.

A close-up of a thickly painted oil painting: a tool wall with neatly hung wrenches, one slot outlined in burnt orange; the small indigo block does not hang on the wall — it sits below, embedded in a network of recessed channels that visibly rearrange themselves around it.
A close-up of a thickly painted oil painting: a tool wall with neatly hung wrenches, one slot outlined in burnt orange; the small indigo block does not hang on the wall — it sits below, embedded in a network of recessed channels that visibly rearrange themselves around it.

A score of 18 did not mean an 18 percent probability of placement, or any other fixed probability. In the first documented version of the Allegheny Family Screening Tool (AFST), it fell within one of the three highest bands. County policy then gave that band a practical consequence: investigation became the prescribed route unless a supervisor approved an override.

The threshold is the clearest place to see what changed. Before 2016, screeners at Allegheny County’s child-welfare hotline mainly assembled information for a supervisor. They could offer an opinion, but the supervisor made the formal decision to investigate a referral or close it without an investigation. With Version 1, screeners were expected to record their own risk and safety assessments, generate a model score, and recommend an outcome for many referrals. Supervisors reviewed those recommendations.

The intervention was therefore larger than the addition of a predictive tool. It changed the division of judgment and attached a rule to one point on the display.

What 18 meant in the model

Version 1 produced two scores for each child. The output that acquired procedural force was the placement-risk score. It estimated the likelihood that a child would be placed outside the home within two years, conditional on the referral first being investigated. It did not determine whether the allegation was true, and it was not a measure of immediate danger. Those distinctions mattered because hotline staff used the language of risk and safety for information arriving in the call itself.

The model was a logistic regression based on linked administrative records about people associated with a referral. It did not use the substance of the allegation. A screener might therefore hear about an acute incident that was entirely absent from the model. In this article, AI refers broadly to a system that derives an assessment from data; nothing in the case depends on a language model or an autonomous agent.

The 1-to-20 display was not a probability scale. Predicted probabilities were divided into twenty bands containing roughly five percent of cases each. Equal steps on the display did not imply equal changes in estimated risk. The county’s boundary at 18 identified the highest three bands, or 15 percent of the distribution.

The county’s published overview dates the launch of Version 1 to August 2016. Version 2 replaced it in December 2018 with a revised algorithm, data sources, and associated policies. Version 1 remains instructive because ordinary regression acquired organizational force through workflow rules, long before the present interest in generative AI.

The number changed who had to disagree

Below 18, the score contributed to a screener’s recommendation. At 18 or above, the system labeled the case a “mandatory screen-in.” Only a supervisor could keep such a referral from proceeding. The result was a default with an exception, not an automated disposition.

Early operating data make that distinction tangible. Supervisors overrode almost one quarter of mandatory screen-ins, and override rates varied between supervisors. Screening outcomes were more closely associated with the workers’ own risk-and-safety assessments than with the AFST score. According to the research team’s case study, staff were not simply rubber-stamping the model.

The overrides did not restore the former process. Once policy made investigation the expected outcome, departing from it required attention, disagreement, and the correct authority. The exception was therefore part of the system’s influence, and variation between supervisors became an implementation issue worth examining.

This is also where evaluation has to follow the authority map rather than the model alone. It needs to identify the claim made by the output, whether that output supplies information, sets a default, or starts an action, and who may depart from it. Measures must cover missed cases, unequal effects, and the burden placed on each role as well as model performance.

The county’s process evaluation documents the broader redistribution of judgment. Adoption alone would be a poor measure of this intervention: it cannot show whether the new arrangement improved decisions or merely moved work between roles.

What the evidence cannot separate

Because the county altered roles and rules at the same time, the deployment was not a controlled test of the score’s effect. Later changes in screening could have come from the prediction, the revised responsibilities, supervisor practice, or their interaction. The case provides rich evidence about implementation but weak evidence for attributing an outcome to the model alone.

The validation history deepens that uncertainty. The developers later reported two mistakes that had made initial performance estimates too optimistic. Repeated referrals for the same child, and records for siblings on the same referral, could appear on opposite sides of the training-test split. Feature selection had also used the full dataset before validation. A corrected procedure changed the performance estimates, though not the deployed model.

Even the revised evaluation faced a “selective labels” problem. Training and testing covered referrals that had been investigated, while the tool was meant to inform selection from all referrals. Comparable later outcomes were not available for many screened-out cases. Fairness could not be reduced to a metric either. Public agencies tend to hold more records about poor families and groups with greater exposure to government services; professional judgment carries biases of its own.

The record therefore leaves the central causal question open. It does not tell us how much of any later change came from the prediction, the new distribution of responsibility, the default at 18, or the interaction among them.

Oliver Wrede writes and teaches on interface design, knowledge systems, and the architecture of intelligence in organizations. He is interested in how humans, institutions, and machines reason together — and how design shapes the quality of that reasoning.

More from Oliver Wrede