What happened
It was reported that Allegheny County, which includes Pittsburgh, used a tool called the Allegheny Family Screening Tool to help its staff decide which reports of possible child neglect to investigate. When a call came in, the tool gave the family a risk score from one to twenty, meant to predict the chance that a child would be placed in foster care in the two years after the report, drawing on government records that included the family’s child welfare history, birth records, use of Medicaid and other benefits, mental health and substance use, jail and probation, and diagnoses of disability. The score was one input for the human screeners, not the whole decision.
It was reported that in 2022 an investigation by the Associated Press raised concerns about the tool. It reported that the tool did not use race directly but leaned on data tied to poverty, disability and the use of public benefits that critics said could stand in for it, that families could see little of how they had been scored, and that researchers who studied the tool found it could flag some groups, including Black families, more often. The US Justice Department later looked at whether the tool discriminated against people with disabilities, including parents with mental health conditions. Allegheny County defended the tool, saying the disability data was predictive of the outcomes it measured and that parents with disabilities might also need extra support, and it revised the tool over time.
What an auditable version would have shown
A family investigated after a report had little way to see what the score was, what data it drew on, or how much it had shaped the decision to knock on their door. An auditable version keeps a record for each report showing the score, the data behind it, and whether a human weighed it and how, and makes it reachable to the family, so a score built on a wrong record, or on data the family disputes, is something they can see and contest. It would also show the wider pattern: how the scores, and the investigations that follow, land across race, disability and income, so a system that falls harder on some families is something the county can measure, not something an outside investigation has to piece together after the fact.
Where the gap was
A score shaped whether a family reported for neglect was investigated, and the family could see little of the score or the data behind it. A ConductRecord keeps each score, the data it used and the human decision that followed, so a family can be shown why they were investigated and can challenge a wrong record. A MetricRecord counts how the scores, and the investigations that follow, land across race, disability and income, so a system that weighs against disabled or poor families shows up as a number the county can see and a regulator can check. The county said the disability data was predictive. Whether using it was fair to the families it counted against is something a record of the outcomes across groups would show.
What governance should have looked like
Where a score shapes a decision as serious as whether to investigate a family for child neglect, the family should be able to see the score and the data behind it and to challenge it, a person should make and record the decision, and the effect of the score across race, disability and income should be measured and open to review. Best practice would be for the county to keep each score, its basis and the human decision, to give a family a real account of why it was investigated, and to measure across groups whether the scoring falls evenly. The county could say the data predicted its chosen outcome. Whether that outcome, and the data standing in for it, treated disabled and poor families fairly is the question the measurement across groups is meant to answer.
Failure Pattern: an algorithm scored families reported for possible neglect on data including disability and poverty to shape which were investigated, and a family could not see the score or the data behind it or readily challenge how they had been judged.
Governance Principle: where a score shapes a child protection decision about a family, the family must be able to see the score and the data behind it and challenge it, a person must make and record the decision, and the score’s effect across race, disability and income must be measured and open to review.
The reference implementation of ConductRecord and MetricRecord is open source. It lives at github.com/saffronandindia/headlights-oss, Apache 2.0 licensed and free to install. The repository is public now.
Sources
- DOJ examining AI screening tool used by Pennsylvania child welfare agency (PBS NewsHour, on the AP investigation)
- Child welfare algorithm used by Allegheny County DHS faces Justice Department scrutiny (WESA)
- CPS is using potentially biased AI screening tools (Popular Science)