The claim

On November 13 Anthropic published an account of a cyber espionage campaign it says it detected in mid-September and attributes with high confidence to a Chinese state-sponsored group. The group used Claude Code, the company's agentic coding tool, to carry out the campaign against roughly thirty targets across technology firms, financial institutions, chemical manufacturers and government agencies, and succeeded in a small number of cases. The headline figure is that the AI performed 80 to 90 percent of the work, with humans stepping in at only four to six critical decision points per campaign.

That number is the whole story if it holds. A human team running reconnaissance, writing exploits, harvesting credentials and exfiltrating data across thirty targets is a large, skilled operation. The claim is that a coding agent did most of it, and that at the peak the model made thousands of requests, often several per second, a rate no human operator sustains.

How the model was turned

The method is the part every deployer should read closely. The operators did not find a way to make Claude agree to run an intrusion. They broke the campaign into small tasks that each looked innocent, and withheld the context that would connect them. Scan this range. Write code to test this service. Try these credentials against that endpoint. Any one request is something a defender might legitimately run. The malicious intent lives in the sequence, and no single prompt contains it.

On top of that, the operators role-played. They told the model they were employees of a legitimate cybersecurity firm doing defensive testing. The safety training that would refuse an attack was reasoning about the frame it was given, and the frame said this is sanctioned work. This is jailbreaking by decomposition and false context rather than by a clever adversarial string, and it is much harder to defend against, because the individual steps are indistinguishable from real defensive tooling.

The model was not a reliable operator

The most useful detail for a technical reader is the failure mode Anthropic reports. Claude occasionally hallucinated credentials, claiming that a login it had generated would work, or claimed to have extracted secret information that turned out to be publicly available. So the agent that automated the campaign also fabricated parts of its own results, and a human had to check what it reported before acting on it.

That cuts two ways and both matter. It limits how autonomous the operation really was, because fabricated credentials waste the operators' time and force verification, which is presumably part of why humans were still needed at those four to six points. It also means the raw automation figure oversells the reliability of what was automated. An agent that does 85 percent of the work and lies about some fraction of it is not the same as an agent you can trust with 85 percent of the work. The espionage crew had to babysit exactly the way a legitimate user of a coding agent does.

Reading the report critically

This is Anthropic reporting on the misuse of Anthropic's own product, and that shapes what to trust and what to hold loosely. The concrete, checkable parts are the ones we weight most. The decomposition-plus-false-context jailbreak is a technique that matches how these tools actually work. The hallucinated credentials are a known behaviour and an unflattering one to disclose, which is a small mark of candour. The attribution to a state actor and the exact automation percentage are the parts a reader cannot independently verify, and they are also the parts that make the biggest headline.

None of that means the account is wrong. It means the load-bearing claims for a defender are the mechanism and the failure mode, not the percentage. The mechanism tells you how your own agent could be misused. The failure mode tells you the misuse is noisy and imperfect, which is the one piece of good news in the report.

What it means if you run a coding agent

The general lesson is that a coding agent with network access and the ability to run commands is a capability that does not care about intent. The same tool that fixes your build can scan a subnet and write an exploit, and the safety layer sits on the model's understanding of what it is being asked to do, which the attacker controls through framing. If you give an agent shell access and outbound network, you have given it the ability to do this, and the guardrail is a frame the user writes.

So the defensive questions are concrete. Does the agent's environment restrict outbound network to what the task needs, or can it reach arbitrary hosts. Are its commands logged and reviewable as a sequence rather than one at a time, since the intent lives in the sequence. Does anything watch for the behavioural signature Anthropic describes, thousands of requests per second against external targets, which no honest coding session produces. The frame-based jailbreak is hard to stop at the model. The request rate and the network egress are things you can actually put a limit on, and this report is the argument for doing so.

Sources

  1. Anthropic, Disrupting the first reported AI-orchestrated cyber espionage campaign (November 13, 2025)