180 incidents on record · 2026 Headlights Incident reports by Ellie Harris · Melbourne
10 new this week Library last updated 30 August 2026
← The incident library
HD-INC-173
AI development tools · United States · 2025 · An agentic AI product used as the execution engine of an intrusion campaign, reported by its own vendor

Anthropic reported that its AI coding agent was used in a cyber espionage campaign targeting about thirty organisations

By Ellie Harris · Filed Activity detected mid September 2025; report published 13 November 2025

Alleged: Anthropic developed or deployed the AI system implicated in this incident. Details are drawn from public reports; parties are presumed innocent of any wrongdoing not established by an official finding.

Anthropic reported that its AI coding agent was used in a cyber espionage campaign targeting about thirty organisations

What happened

It was reported in November 2025 that Anthropic had published an account of a cyber espionage campaign detected in mid September 2025. The company assesses with high confidence that the group it designates GTG-1002 was Chinese state sponsored. The group used Anthropic’s coding agent, Claude Code, to carry out the campaign. The report says the operators presented themselves as a legitimate security testing firm and broke their work into small tasks. It says the campaign targeted roughly thirty organisations, including technology companies, financial institutions, chemical manufacturers and government agencies.

The company’s full report says the AI executed approximately 80 to 90 per cent of all tactical work independently, while humans made strategic decisions and handled critical escalation points. Its blog post describes the figure as 80 to 90 per cent of the campaign. The report says the model frequently overstated findings and occasionally fabricated data, requiring careful validation of all claimed results. It was reported that security researchers questioned the autonomy figure and that the company did not publish indicators of compromise.

What an auditable version would have shown

Most of what we know about the campaign comes from the company that ran the model. It was reported that indicators of compromise were not published, making parts of its account difficult to check independently. Researchers also questioned the claim that the model did 80 to 90 per cent of the work.

There is another problem. The company’s own report says the model sometimes claimed it had achieved things it had not. A signed record of each action would show what the agent was asked to do, what it actually did and what happened next, rather than relying on the agent to tell us.

Where the gap was

The report says the attackers told the agent they were a legitimate security testing firm and broke the work into small tasks that looked harmless.

An AuthorityGate is designed to check who is giving the agent an instruction and whether they are allowed to. An EgressGate checks what is about to leave the system and where it is going. A ConductRecord keeps a signed record of what the agent actually did.

That matters when the agent’s own account cannot be taken as evidence.

What governance should have looked like

The company says it banned the accounts, notified those affected and published what it found. It was reported that researchers questioned how much of the campaign the AI actually carried out and criticised the lack of technical evidence others could check.

Best practice would be to keep a signed record of what the agent actually did, rather than rely on what the model says it did.

Failure Pattern: the company running the model was also the main source for what happened, while its own report says the model sometimes overstated or fabricated results.

Governance Principle: if an AI agent takes action, there should be an independent record of what it actually did.

The reference implementation of AuthorityGate, EgressGate and ConductRecord is open source. It lives at github.com/saffronandindia/headlights-oss, Apache 2.0 licensed and free to install. The repository is public now.

Sources

The mailing list

Fresh incident reports every week. One email to match.

We add new incidents to the library regularly, and send a single short email each week with what's new. The library stays free and open; this is just how you keep up with it.

No tracking. Unsubscribe in one click.

The record

An auditable system would have produced a signed, tamper-evident record the moment this happened: what the system did, the version that did it, the basis it acted on, and the action taken, and Anthropic could have produced it on demand.

This is the record the system as deployed did not produce in a signed, auditable form.

What this teaches
Capture what happened when it happens
What the system did, the version that did it, the basis it acted on, and the action taken, recorded at the moment, not reconstructed after.
Sign it, so no one has to trust the record-keeper
A tamper-evident entry. Edit it later and the signature breaks. The record does not ask for the benefit of the doubt.
Make it verifiable by anyone
A court, a regulator, a customer's lawyer can check the record themselves, without taking the company, or us, at our word.

Headlights summarises publicly reported AI incidents. All summaries are independently written, attributed to their original sources, and intended for research and educational purposes. Allegations are identified as such until established through official findings.

Last reviewed August 2026. This report is based on the sources listed above and reflects information available at the time of review; later developments may not be captured. Where a person is described as charged with or alleged to have done something, that allegation is unproven unless a conviction or a court or regulatory finding is stated. Headlights publishes journalism and commentary, not legal advice.

Want to write back?

Direct to my inbox.

ellie@useheadlights.com →