What happened
It was reported that the Wisconsin Department of Public Instruction ran a system called the Dropout Early Warning System, or DEWS, that scored middle-school students on how likely they were to finish high school on time. Since 2012 it had, year after year, produced a prediction for hundreds of thousands of children in grades six to nine, drawing on test scores, discipline records, whether a student qualifies for free or reduced-price lunch, and their race, and it showed the results to their schools. According to an investigation by The Markup published in April 2023, DEWS raised false alarms, predicting that a student would not graduate on time when in fact they did, about 42 percentage points more often for Black students than for white students, and about 18 points more often for Hispanic students. The Markup also reported that when DEWS predicted a student would not graduate, it was wrong about 74 per cent of the time.
It was reported that the department already knew. An internal equity analysis it carried out in 2021 had asked plainly whether DEWS was fair and answered no, finding that the model over-identified white students among those who graduated on time and over-identified Black, Hispanic and other students of colour among those who did not. The Markup reported that the department did not pass this finding on to the schools that used DEWS, and did not change the system, in the nearly two years between reaching it and being asked about it. A department spokesperson defended the system and, asked about the disparities, told The Markup that the questioner had a fundamental misunderstanding of how it works. After the investigation, in October 2023, the department removed the dropout-warning data from the dashboards schools used and said it was reviewing the future of the system.
What an auditable version would have shown
DEWS gave every child a prediction, and the schools acting on it had no way to see how often it was right, or which children it was getting wrong. An auditable version keeps a record of each prediction next to what actually happened to that student, so the rate at which the system is wrong, and how that rate differs between groups of children, is a figure the department holds and can be asked for. The department did work that figure out once, in 2021. What an auditable system does is keep it in front of the people using the tool, rather than in an internal review they were never shown.
Where the gap was
A prediction about a child was handed to a school as guidance, and nothing told the school how often the guidance was wrong, or that it was wrong far more often for Black and Hispanic students. A MetricRecord counts how the predictions perform against what actually happens, broken down by group, so a pattern of false alarms is a figure the people using the tool can see rather than one sitting in a single internal review. A ConductRecord keeps each prediction and what it rested on, so a child wrongly flagged can be found and the flag looked at again. The department had measured the disparity and found the system unfair. The schools using it were shown neither the measurement nor the finding, and the tool kept running.
What governance should have looked like
Where a system predicts something about a child and shows it to their school, its accuracy should be measured against real outcomes, broken down by the groups it affects, and that measurement should be put in front of the people who rely on it rather than held internally. A child marked as a risk should also be able to have that mark reviewed. Wisconsin’s department had the measurement, and its own finding that the system was unfair. What it did not do was act on either or tell the schools, and the predictions went on being shown to teachers as though they could be trusted.
Failure Pattern: a predictive score about children was shown to their schools as guidance, with no measure of how often it was wrong or how much more often it was wrong for some groups, and a finding that it was unfair was kept internal.
Governance Principle: where a system predicts something about a person and others act on it, its accuracy must be measured against real outcomes, broken down by the groups it affects, and that measure must be shown to the people who rely on it.
The reference implementation of MetricRecord and ConductRecord is open source. It lives at github.com/saffronandindia/headlights-oss, Apache 2.0 licensed and free to install. The repository is public now.
Sources
- False Alarm: How Wisconsin Uses Race and Income to Label Students ‘High Risk’ (The Markup)
- We’re Not Living a ‘Predicted’ Life: student perspectives on Wisconsin’s dropout algorithm (The Markup, Dec 2023 follow-up)
- Wisconsin uses race and income to label students ‘high risk’ (Wisconsin Watch)