180 incidents on record · 2026 Headlights Incident reports by Ellie Harris · Melbourne
10 new this week Library last updated 30 August 2026
← The incident library
HD-INC-139
Consumer AI · Germany · 2023 · Unvetted training data

Researchers found suspected child-abuse images inside the open dataset used to train popular AI image generators

By Ellie Harris · Filed Stanford Internet Observatory report 20 December 2023

Alleged: LAION e.V. (dataset); models including Stable Diffusion (Stability AI) developed or deployed the AI system implicated in this incident. Details are drawn from public reports; parties are presumed innocent of any wrongdoing not established by an official finding.

Researchers found suspected child-abuse images inside the open dataset used to train popular AI image generators

What happened

In December 2023, researchers at the Stanford Internet Observatory examined LAION-5B, a public dataset containing more than five billion image-text pairs that had been used to train AI image generators, including Stable Diffusion.

What they found was deeply concerning.

The researchers identified 3,226 suspected instances of child sexual abuse material, with about 1,008 later validated by external child safety organisations. They used automated detection methods, including matching against databases of known illegal material, and worked with specialist organisations to confirm and report what they found. The process was designed to minimise direct exposure to the material.

LAION said it had a strict policy against this content and took the dataset offline while it investigated. It later released a cleaned version called Re-LAION-5B.

The incident highlighted a much bigger problem. The dataset was so large that nobody had fully checked what was inside before it became training data for AI models. Researchers warned that if prohibited material enters a training dataset, it may influence the behaviour of models trained on it, even if the impact on any individual model cannot be measured afterwards.

What an auditable version would have shown

An auditable training dataset should tell the story of where each piece of data came from. It should show what checks were carried out before the data was added, what screening tools were used, what was flagged, what was removed, and who approved the dataset for training. That way, if someone later asks what a model was trained on, the answer already exists. It doesn’t have to be discovered by researchers months or years later.

Where the gap was

The problem wasn’t simply that prohibited material existed. It was that the dataset had grown to a size where nobody could confidently say what was inside it. Collecting billions of images from the open internet is relatively easy. Proving each one of those images is suitable for training is much harder. The scale that made the dataset valuable also made it difficult to verify.

What governance should have looked like

Training data should be treated like any other critical asset. Organisations should know where it came from, what checks were performed before it was included, and who approved it for use. Each new item should be screened against known databases of prohibited material before it becomes part of a training dataset, with a clear record of what was blocked, what was reviewed, and why. The goal isn’t to slow innovation. It’s to make sure organisations can answer a simple question with evidence: what exactly did this model learn from?

A ConductRecord of where each part of a dataset came from, and a ConstraintGate that screens incoming material against known-bad sources and blocks what it matches, would mean prohibited content is caught at the door rather than found later inside a shipped model.

The reference implementation of ConductRecord and ConstraintGate is open source. It lives at github.com/saffronandindia/headlights-oss, Apache 2.0 licensed, free for any company to install. The repository is public now.

Sources

The mailing list

Fresh incident reports every week. One email to match.

We add new incidents to the library regularly, and send a single short email each week with what's new. The library stays free and open; this is just how you keep up with it.

No tracking. Unsubscribe in one click.

The record

An auditable system would have produced a signed, tamper-evident record the moment this happened: what the system did, the version that did it, the basis it acted on, and the action taken, and LAION e.V. (dataset); models including Stable Diffusion (Stability AI) could have produced it on demand.

This is the record the system as deployed did not produce in a signed, auditable form.

What this teaches
Capture what happened when it happens
What the system did, the version that did it, the basis it acted on, and the action taken, recorded at the moment, not reconstructed after.
Sign it, so no one has to trust the record-keeper
A tamper-evident entry. Edit it later and the signature breaks. The record does not ask for the benefit of the doubt.
Make it verifiable by anyone
A court, a regulator, a customer's lawyer can check the record themselves, without taking the company, or us, at our word.

Headlights summarises publicly reported AI incidents. All summaries are independently written, attributed to their original sources, and intended for research and educational purposes. Allegations are identified as such until established through official findings.

This report is based on the sources listed above and reflects information available at the time of review; later developments may not be captured. Where a person is described as charged with or alleged to have done something, that allegation is unproven unless a conviction or a court or regulatory finding is stated. Headlights publishes journalism and commentary, not legal advice.

Want to write back?

Direct to my inbox.

ellie@useheadlights.com →