180 incidents on record · 2026 Headlights Incident reports by Ellie Harris · Melbourne
10 new this week Library last updated 30 August 2026
← The incident library
HD-INC-134
Consumer AI · United States · 2025 · Inadequate child-safety guardrails

Reuters found an internal Meta policy that had allowed its AI chatbots to hold romantic or sensual conversations with children, and Meta said the passages were a mistake and had been removed

By Ellie Harris · Filed Reuters reported 14 August 2025

Alleged: Meta Platforms, Inc. developed or deployed the AI system implicated in this incident. Details are drawn from public reports; parties are presumed innocent of any wrongdoing not established by an official finding.

Reuters found an internal Meta policy that had allowed its AI chatbots to hold romantic or sensual conversations with children, and Meta said the passages were a mistake and had been removed

What happened

It was reported that in August 2025 Reuters obtained an internal Meta document, running to about 200 pages and titled GenAI: Content Risk Standards, that set out what the company’s AI chatbots were allowed to produce. Reuters reported that the standards had, among their examples, permitted a chatbot to engage a child in conversation that was romantic or sensual, while stating it was unacceptable to describe sexual actions to a child when roleplaying. The reporting said the same document allowed other things as well, including stating false information where it was labelled as untrue, and, under set conditions, content that demeaned people on the basis of protected characteristics.

Reuters reported that Meta confirmed the document was genuine. A spokesperson said the passages about children were erroneous and inconsistent with the company’s policies, that they should not have been there, and that they had been removed. Meta said it does not permit sexualised interactions with children, and noted its chatbots are open to users aged 13 and over. After the report, Senator Josh Hawley opened an inquiry into Meta’s AI policies, other members of Congress raised concerns, and child-safety advocates said the assurances did not go far enough and asked Meta to publish its current rules. Meta later said it would add temporary limits on how its AI chatbots engage with teenagers.

What an auditable version would have shown

Every rule an AI chatbot follows was written by a person and approved by a person. So why don’t we treat those rules like every other critical record? We should know who wrote them, who approved them, when they changed, and why. Some rules should be locked so they can’t be quietly altered later. If someone weakened an important safeguard, there should be a clear record showing exactly who made that decision, instead of us discovering it years later after a whistleblower or journalist exposes it.

Where the gap was

The problem wasn’t that there wasn’t a rule. The problem was that a rule which should have been absolute could still be weakened. A safeguard as simple as no romantic or sexual content involving a child should never be treated like an ordinary instruction that can be rewritten or overridden. It should be permanent, visible and easy to verify. Instead, it sat inside a long policy document until it was exposed through a leak.

What governance should have looked like

For products used by children, the guardrails aren’t an extra feature, they are the product.

A chatbot designed to feel like a friend should have child safety built in as a non-negotiable limit. No matter what a user asks, what character the chatbot is playing, or what later changes are made to the system, that safeguard should always hold.

Those critical rules should have a named author, a named approver, and a record of every change. They should also be tested before the product is released, with evidence that they cannot be bypassed. Parents shouldn’t have to trust that these protections exist. Companies should be able to prove they do.

The reference implementation of ConstraintGate, PersonaGuard and ConductRecord is open source. It lives at github.com/saffronandindia/headlights-oss, Apache 2.0 licensed, free for any company to install. The repository is public now.

Sources

The mailing list

Fresh incident reports every week. One email to match.

We add new incidents to the library regularly, and send a single short email each week with what's new. The library stays free and open; this is just how you keep up with it.

No tracking. Unsubscribe in one click.

The record

An auditable system would have produced a signed, tamper-evident record the moment this happened: what the system did, the version that did it, the basis it acted on, and the action taken, and Meta Platforms, Inc. could have produced it on demand.

This is the record the system as deployed did not produce in a signed, auditable form.

What this teaches
Capture what happened when it happens
What the system did, the version that did it, the basis it acted on, and the action taken, recorded at the moment, not reconstructed after.
Sign it, so no one has to trust the record-keeper
A tamper-evident entry. Edit it later and the signature breaks. The record does not ask for the benefit of the doubt.
Make it verifiable by anyone
A court, a regulator, a customer's lawyer can check the record themselves, without taking the company, or us, at our word.

Headlights summarises publicly reported AI incidents. All summaries are independently written, attributed to their original sources, and intended for research and educational purposes. Allegations are identified as such until established through official findings.

This report is based on the sources listed above and reflects information available at the time of review; later developments may not be captured. Where a person is described as charged with or alleged to have done something, that allegation is unproven unless a conviction or a court or regulatory finding is stated. Headlights publishes journalism and commentary, not legal advice.

Want to write back?

Direct to my inbox.

ellie@useheadlights.com →