180 incidents on record · 2026 Headlights Incident reports by Ellie Harris · Melbourne
10 new this week Library last updated 30 August 2026
← The incident library
HD-INC-121
Healthcare · United States · 2019 · Cost used as a biased proxy for health need

A widely used US healthcare algorithm judged how sick patients were by what had been spent on their care, so patients the system had historically spent less on were rated healthier and passed over for extra help, and researchers found the bias cut those referrals by more than half

By Ellie Harris · Filed Study published October 2019

Alleged: Optum (UnitedHealth Group) developed or deployed the AI system implicated in this incident. Details are drawn from public reports; parties are presumed innocent of any wrongdoing not established by an official finding.

A widely used US healthcare algorithm judged how sick patients were by what had been spent on their care, so patients the system had historically spent less on were rated healthier and passed over for extra help, and researchers found the bias cut those referrals by more than half

What happened

It was reported that hospitals across the United States used a risk-scoring tool, made by Optum, part of UnitedHealth Group, to decide which patients to enrol in programmes that give extra care to the sickest. The tool judged how sick a patient was by predicting how much they would cost the health system, treating that cost as a stand-in for need. Researchers who studied it at a large hospital, in a 2019 paper in the journal Science, found a hidden bias in that substitution. Because the health system had long spent less caring for Black patients than for equally sick white patients, the tool read the lower spending as better health and marked them as healthier than they were.

It was reported that the effect was large. At any given score, the researchers found, those patients were in fact sicker, carrying more chronic illness than others rated the same, and rebuilding the tool to predict illness rather than cost would have raised their share of the places for extra care from about 18 percent to almost half. Tools of this kind, the team estimated, were being applied to around 200 million people a year. Optum said the model was only one of many things a hospital was meant to weigh, and called the researchers’ conclusions misleading. The tool never used race. The bias sat in the choice to let past spending stand for need, on patients the system had spent less to treat.

What an auditable version would have shown

The makers of the tool and the hospitals using it could see the scores it produced, but not what those scores did across different groups of patients. What the researchers had to reconstruct from the data, how sick Black and white patients were at the same score, and how many of each were referred for extra care, is exactly what a system built to be checked would already hold. An auditable version keeps, alongside each score, the outcome it was meant to predict and the actual health of the patient, so the gap between what the tool measured, cost, and what it was standing in for, need, is a figure the health system can see rather than one an outside study has to uncover years later.

Where the gap was

A tool predicted cost and was used as though it predicted need, and what that substitution did across different groups of patients was surfaced by outside researchers rather than visible in the tool’s own output. A MetricRecord keeps the bigger picture: whether the scores, and the care that came with them, matched how sick each group of patients really was, so a tool that quietly under-serves one group shows up as a number the health system already holds, not a finding it has to learn from a journal. A ConductRecord keeps each scoring decision and the data behind it, so a patient wrongly passed over can be traced. The tool never used race as an input. A record makes the effect visible, and the effect is where the harm was.

What governance should have looked like

Where a score decides who receives extra care, the thing it predicts has to be the thing that matters, and where a proxy is used in its place, the effect of that proxy across the people it sorts has to be measured. Best practice would be for a health system to record, for each patient, the score, what it was predicting, and the care that followed, and to measure across groups whether the score tracks real health need or something else, like spending, that runs alongside it. The researchers could show the bias because they had the underlying data. A health system that measured this for itself would not need a study to tell it that its sickest patients were being missed.

Failure Pattern: a score meant to find the patients most in need of care predicted their cost instead, so a group on whom less had historically been spent was rated healthier and referred less often, and the disparity was surfaced by outside researchers rather than shown by the tool itself.

Governance Principle: where a score decides who receives care or a benefit, it must predict the thing that matters rather than a proxy that tracks it unevenly, and its effect across the groups it sorts must be measured against the real outcome.

The reference implementation of MetricRecord and ConductRecord is open source. It lives at github.com/saffronandindia/headlights-oss, Apache 2.0 licensed and free to install. The repository is public now.

Sources

The mailing list

Fresh incident reports every week. One email to match.

We add new incidents to the library regularly, and send a single short email each week with what's new. The library stays free and open; this is just how you keep up with it.

No tracking. Unsubscribe in one click.

The record

An auditable system would have produced a signed, tamper-evident record the moment this happened: what the system did, the version that did it, the basis it acted on, and the action taken, and Optum (UnitedHealth Group) could have produced it on demand.

This is the record the system as deployed did not produce in a signed, auditable form.

What this teaches
Capture what happened when it happens
What the system did, the version that did it, the basis it acted on, and the action taken, recorded at the moment, not reconstructed after.
Sign it, so no one has to trust the record-keeper
A tamper-evident entry. Edit it later and the signature breaks. The record does not ask for the benefit of the doubt.
Make it verifiable by anyone
A court, a regulator, a customer's lawyer can check the record themselves, without taking the company, or us, at our word.

Headlights summarises publicly reported AI incidents. All summaries are independently written, attributed to their original sources, and intended for research and educational purposes. Allegations are identified as such until established through official findings.

This report is based on the sources listed above and reflects information available at the time of review; later developments may not be captured. Where a person is described as charged with or alleged to have done something, that allegation is unproven unless a conviction or a court or regulatory finding is stated. Headlights publishes journalism and commentary, not legal advice.

Want to write back?

Direct to my inbox.

ellie@useheadlights.com →