What happened
It was reported that hospitals across the United States used a risk-scoring tool, made by Optum, part of UnitedHealth Group, to decide which patients to enrol in programmes that give extra care to the sickest. The tool judged how sick a patient was by predicting how much they would cost the health system, treating that cost as a stand-in for need. Researchers who studied it at a large hospital, in a 2019 paper in the journal Science, found a hidden bias in that substitution. Because the health system had long spent less caring for Black patients than for equally sick white patients, the tool read the lower spending as better health and marked them as healthier than they were.
It was reported that the effect was large. At any given score, the researchers found, those patients were in fact sicker, carrying more chronic illness than others rated the same, and rebuilding the tool to predict illness rather than cost would have raised their share of the places for extra care from about 18 percent to almost half. Tools of this kind, the team estimated, were being applied to around 200 million people a year. Optum said the model was only one of many things a hospital was meant to weigh, and called the researchers’ conclusions misleading. The tool never used race. The bias sat in the choice to let past spending stand for need, on patients the system had spent less to treat.
What an auditable version would have shown
The makers of the tool and the hospitals using it could see the scores it produced, but not what those scores did across different groups of patients. What the researchers had to reconstruct from the data, how sick Black and white patients were at the same score, and how many of each were referred for extra care, is exactly what a system built to be checked would already hold. An auditable version keeps, alongside each score, the outcome it was meant to predict and the actual health of the patient, so the gap between what the tool measured, cost, and what it was standing in for, need, is a figure the health system can see rather than one an outside study has to uncover years later.
Where the gap was
A tool predicted cost and was used as though it predicted need, and what that substitution did across different groups of patients was surfaced by outside researchers rather than visible in the tool’s own output. A MetricRecord keeps the bigger picture: whether the scores, and the care that came with them, matched how sick each group of patients really was, so a tool that quietly under-serves one group shows up as a number the health system already holds, not a finding it has to learn from a journal. A ConductRecord keeps each scoring decision and the data behind it, so a patient wrongly passed over can be traced. The tool never used race as an input. A record makes the effect visible, and the effect is where the harm was.
What governance should have looked like
Where a score decides who receives extra care, the thing it predicts has to be the thing that matters, and where a proxy is used in its place, the effect of that proxy across the people it sorts has to be measured. Best practice would be for a health system to record, for each patient, the score, what it was predicting, and the care that followed, and to measure across groups whether the score tracks real health need or something else, like spending, that runs alongside it. The researchers could show the bias because they had the underlying data. A health system that measured this for itself would not need a study to tell it that its sickest patients were being missed.
Failure Pattern: a score meant to find the patients most in need of care predicted their cost instead, so a group on whom less had historically been spent was rated healthier and referred less often, and the disparity was surfaced by outside researchers rather than shown by the tool itself.
Governance Principle: where a score decides who receives care or a benefit, it must predict the thing that matters rather than a proxy that tracks it unevenly, and its effect across the groups it sorts must be measured against the real outcome.
The reference implementation of MetricRecord and ConductRecord is open source. It lives at github.com/saffronandindia/headlights-oss, Apache 2.0 licensed and free to install. The repository is public now.
Sources
- Dissecting racial bias in an algorithm used to manage the health of populations (Obermeyer et al., Science, 2019)
- Millions of black people affected by racial bias in health-care algorithms (Nature)
- Study finds racial bias in Optum algorithm (Healthcare Finance News)