On September 11, Anthropic released a 154-page report detailing how state-linked actors used Claude between December 2025 and August 2026. Five operations. Four countries. The one that reads like a thriller involves an Iran-linked group that turned a chatbot into a naval intelligence desk.
What the Iranian operation looked like
The operator did not ask Claude to plan an attack. It asked Claude to organize publicly available information: ship transponder identifiers, aircraft transponder codes, commercial satellite imagery, and crew names scraped from captions on public military photographs. Individually, each request was unremarkable. A question about a ship identifier. A request to parse a satellite image service's API. A query about the naming convention for Navy personnel records.
Assembled, the pieces formed a targeting handbook for U.S. naval forces in the Middle East. The operator also directed Claude to compile vulnerability research on shipboard systems — known cybersecurity flaws in maritime VSAT terminals, Cisco communications equipment, and industrial control products used aboard vessels.
Every input was technically benign. The output was not. Anthropic's report describes this as deliberate fragmentation: breaking a prohibited goal into individually permissible steps so no single request triggers a safety filter. If you have used a chatbot to assemble scattered information into a coherent brief, you have used the same workflow. The difference is intent, and intent is exactly what a language model cannot see.
The other four operations
The Iran case is the most detailed, but the report documents four more:
Yemen (Houthi-linked): Users in northern Yemen attempted to use Claude to develop software for advanced missile systems, including guidance code for ballistic missiles. Anthropic reports it blocked some requests, but others circumvented safeguards before accounts were banned.
China (government-aligned): Operations leveraged Claude for surveillance and profiling, including targeting Uyghurs and religious figures. The report does not specify what profiling outputs looked like, but the pattern is consistent with the kind of dossier-building that the Iranian operation made explicit.
Russia (state-linked): Actors attempted to use Claude to develop autonomous drone swarms using operational data from the war in Ukraine. The report says these were detected and disrupted.
Across all five: Anthropic says every operation was disrupted, accounts were banned, and findings were shared with relevant authorities. The company describes strengthened detection systems as a result.
What the safety systems caught — and what they did not
The honest reading of this report is that Anthropic's filters worked at the request level and failed at the campaign level. No single prompt in the Iranian operation asked for something explicitly dangerous. The danger emerged from aggregation — assembling individually harmless outputs into something with military utility. That is a harder problem than filtering for "how to build a bomb," and it is the problem that actually matters.
Anthropic deserves credit for publishing this. Most companies that discover state-linked misuse of their products describe it in a paragraph, not 154 pages. The report names countries, describes tradecraft, and acknowledges that some requests slipped through before detection. That level of specificity is useful. It is also, inevitably, a case for Anthropic's own policy positions: the report dropped one day before Dario Amodei published his essay calling for pacing the frontier, and the timing is not accidental.
What this means for you
If you use Claude or any other large language model, you are using the same product that an Iranian intelligence operation used. The interface is the same. The capabilities are the same. The difference is what you ask for and why.
That is not an argument against using the tools. It is an argument for understanding that safety in AI is not a binary — safe or unsafe, filtered or unfiltered. It is a spectrum defined by context, intent, and aggregation. A model that can help you compile research for a presentation can help someone else compile research for a targeting handbook. The question is not whether the tool can be misused. The report proves it can. The question is whether detection systems can improve fast enough to catch campaign-level misuse before it produces real-world harm.
On that question, Anthropic's own report is not reassuring. It describes every operation as disrupted — but the operations existed long enough to produce the outputs described. The Iranian operator got a targeting handbook. The Houthi-linked users got some missile guidance code before their accounts were banned. Disrupted after the fact is not the same as prevented.
That gap — between what filters catch at the prompt level and what campaigns accomplish at the aggregate level — is the unsolved problem in AI safety. And it is the reason Anthropic's CEO is now asking the entire industry to slow down.



