What happened
It was reported that in August 2025 Reuters obtained an internal Meta document, running to about 200 pages and titled GenAI: Content Risk Standards, that set out what the company’s AI chatbots were allowed to produce. Reuters reported that the standards had, among their examples, permitted a chatbot to engage a child in conversation that was romantic or sensual, while stating it was unacceptable to describe sexual actions to a child when roleplaying. The reporting said the same document allowed other things as well, including stating false information where it was labelled as untrue, and, under set conditions, content that demeaned people on the basis of protected characteristics.
Reuters reported that Meta confirmed the document was genuine. A spokesperson said the passages about children were erroneous and inconsistent with the company’s policies, that they should not have been there, and that they had been removed. Meta said it does not permit sexualised interactions with children, and noted its chatbots are open to users aged 13 and over. After the report, Senator Josh Hawley opened an inquiry into Meta’s AI policies, other members of Congress raised concerns, and child-safety advocates said the assurances did not go far enough and asked Meta to publish its current rules. Meta later said it would add temporary limits on how its AI chatbots engage with teenagers.
What an auditable version would have shown
Every rule an AI chatbot follows was written by a person and approved by a person. So why don’t we treat those rules like every other critical record? We should know who wrote them, who approved them, when they changed, and why. Some rules should be locked so they can’t be quietly altered later. If someone weakened an important safeguard, there should be a clear record showing exactly who made that decision, instead of us discovering it years later after a whistleblower or journalist exposes it.
Where the gap was
The problem wasn’t that there wasn’t a rule. The problem was that a rule which should have been absolute could still be weakened. A safeguard as simple as no romantic or sexual content involving a child should never be treated like an ordinary instruction that can be rewritten or overridden. It should be permanent, visible and easy to verify. Instead, it sat inside a long policy document until it was exposed through a leak.
What governance should have looked like
For products used by children, the guardrails aren’t an extra feature, they are the product.
A chatbot designed to feel like a friend should have child safety built in as a non-negotiable limit. No matter what a user asks, what character the chatbot is playing, or what later changes are made to the system, that safeguard should always hold.
Those critical rules should have a named author, a named approver, and a record of every change. They should also be tested before the product is released, with evidence that they cannot be bypassed. Parents shouldn’t have to trust that these protections exist. Companies should be able to prove they do.
The reference implementation of ConstraintGate, PersonaGuard and ConductRecord is open source. It lives at github.com/saffronandindia/headlights-oss, Apache 2.0 licensed, free for any company to install. The repository is public now.
Sources
- Leaked Meta AI rules show chatbots were allowed to have romantic chats with kids (TechCrunch)
- Meta let chatbots engage in inappropriate conversations with minors, Reuters reports (Business & Human Rights Resource Centre)