180 incidents on record · 2026 Headlights Incident reports by Ellie Harris · Melbourne
10 new this week Library last updated 30 August 2026
← The incident library
HD-INC-146
Consumer AI · United States · 2018 · Inaccurate and biased face matching

A civil liberties test ran all of Congress through Amazon's face tool, and it matched 28 of them to mugshots

By Ellie Harris · Filed ACLU test published July 2018

Alleged: Amazon (Rekognition); test run by the ACLU developed or deployed the AI system implicated in this incident. Details are drawn from public reports; parties are presumed innocent of any wrongdoing not established by an official finding.

A civil liberties test ran all of Congress through Amazon's face tool, and it matched 28 of them to mugshots

What happened

It was reported that in July 2018 the American Civil Liberties Union tested Amazon’s face-matching tool, Rekognition. The ACLU reported that it ran photos of all members of the US Congress against 25,000 publicly available arrest photos, using the tool’s default settings and an 80 percent confidence threshold, and that the tool falsely matched 28 members of Congress to police mugshots. The ACLU said nearly 40 percent of those false matches were people of colour, even though they made up about 20 percent of Congress, and that six of those wrongly matched were members of the Congressional Black Caucus. The whole test, it noted, cost 12.33 dollars.

Amazon said the ACLU had used the default 80 percent setting, and that for law enforcement it recommended a higher confidence threshold and human review, while critics pointed out that the default was the setting many customers would use out of the box. In June 2020, amid wider protests over policing, Amazon announced a one-year moratorium on police use of Rekognition, and in 2021 it said the moratorium would continue indefinitely. It continued offering Rekognition for other, non-police image-analysis uses.

What an auditable version would have shown

If you are going to use a tool to identify people, you need to know how often it gets that identification wrong. That means testing the settings customers actually use and breaking the results down by skin tone, age and sex, so a buyer can see the false-match rate rather than just a confidence score attached to one result, and so a match is presented as something to investigate rather than as a fact. Those records would have shown buyers what happened at the default setting, including who was more likely to be falsely matched, before the technology was ever pointed at real people.

Where the gap was

The problem was not simply that Rekognition made mistakes. It was what the person using it could see when it did. A confident-looking match appeared without a clear explanation of how often the system could be wrong, the default setting produced false matches, and nothing in front of the buyer showed that those errors fell more heavily on people of colour. When the consequence of a wrong identification could involve the police, that information matters.

What governance should have looked like

Before a face-matching tool is used to identify people, it should be tested for accuracy and bias using the same settings customers will actually run. A VerificationGate is designed to hold a tool like this against accuracy and fairness tests before it reaches a setting where a wrong match matters, and if it does not pass, it does not move forward. A MetricRecord would keep the evidence behind that decision, including match and error rates broken down by group, so buyers and regulators could see the numbers rather than relying on a general claim of accuracy. The question should be answered before deployment: how often is this system wrong, and who is it getting wrong?

The reference implementation of VerificationGate and MetricRecord is open source. It lives at github.com/saffronandindia/headlights-oss, Apache 2.0 licensed, free for any company to install. The repository is public now.

Sources

The mailing list

Fresh incident reports every week. One email to match.

We add new incidents to the library regularly, and send a single short email each week with what's new. The library stays free and open; this is just how you keep up with it.

No tracking. Unsubscribe in one click.

The record

An auditable system would have produced a signed, tamper-evident record the moment this happened: what the system did, the version that did it, the basis it acted on, and the action taken, and Amazon (Rekognition); test run by the ACLU could have produced it on demand.

This is the record the system as deployed did not produce in a signed, auditable form.

What this teaches
Capture what happened when it happens
What the system did, the version that did it, the basis it acted on, and the action taken, recorded at the moment, not reconstructed after.
Sign it, so no one has to trust the record-keeper
A tamper-evident entry. Edit it later and the signature breaks. The record does not ask for the benefit of the doubt.
Make it verifiable by anyone
A court, a regulator, a customer's lawyer can check the record themselves, without taking the company, or us, at our word.

Headlights summarises publicly reported AI incidents. All summaries are independently written, attributed to their original sources, and intended for research and educational purposes. Allegations are identified as such until established through official findings.

This report is based on the sources listed above and reflects information available at the time of review; later developments may not be captured. Where a person is described as charged with or alleged to have done something, that allegation is unproven unless a conviction or a court or regulatory finding is stated. Headlights publishes journalism and commentary, not legal advice.

Want to write back?

Direct to my inbox.

ellie@useheadlights.com →