Anthropic published Detecting and countering misuse of AI: September 2026 on September 10, 2026, roughly 36,000 words covering “activity we disrupted between December 2025 and August 2026 across seven harm areas”. It is the company’s fourth threat report, after March, August and November 2025, and by far the widest: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons and illicit distillation, each with named case studies.
Yesterday we wrote up Anthropic’s alignment assessment, the report about what its own models did when a sandbox wall was missing. This report is the other side: what people did with Claude on purpose. This page keeps to Anthropic’s document. Every quotation is from it, every number is the report’s own, and the parts that matter most to anyone running agents on real systems, the theft of API keys and sessions, get the most room.
On this page · 10 sectionsOpen
- The two trends Anthropic leads with
- The cyber cases
- The AI supply chain is now the target
- Distillation: seven labs, one at 151 million exchanges
- The other five harm areas
- Which models, and what Anthropic changed
- Anthropic’s threat reports so far
- What this means if you run agents on real systems
- What we are watching for
- Sources
Anthropic published ‘Detecting and countering misuse of AI: September 2026’ on September 10, 2026. It covers activity disrupted between December 2025 and August 2026 across seven harm areas and roughly 40 internally tracked groups, its fourth threat report after March, August and November 2025.
The two trends Anthropic leads with: AI has ‘collapsed the labor and tooling gap’ between state operations and individuals, and ‘a majority of the operations described in this report were enabled by AI via direct execution or orchestration’ through multi-agent frameworks, not chatbot questions and answers.
The AI supply chain is the new target. One Russian-speaking actor injected malicious instructions into an AI vendor’s evaluation sandbox, walked away with its production API keys, then hit about 30 AI companies in about four days. The stated goal, a pre-release Claude model, was never reached, and Anthropic says its own systems were not compromised.
Stolen access is a market. Fake ‘discounted Claude’ resellers shipped credential harvesters disguised as Claude Code installers, one seller offered access that was ‘neither cheap nor actually Claude’, and several actors used prompt injection against LiteLLM wrapper services to exfiltrate production keys.
Seven China-based labs ran distillation campaigns against Claude. Alibaba’s is ‘the largest distillation attack we have ever measured’, over 151 million exchanges between May and July 2026. Moonshot and DeepSeek went further and silently relayed their own users’ requests to Claude, then saved the exchanges for training.
Beyond cyber: nine influence operations on six continents, a Mali surveillance platform built to monitor roughly 25 million SIM cards, a dating-app network with over 4,700 AI personas talking to at least 25,000 people, six conventional-weapons cases and five biological ones.
Every misuse case ran on Claude Haiku, Sonnet or Opus. ‘None of the misuse cases involved the use of Claude Fable or Mythos-class models, with the exception of one illicit distillation case.’ Anthropic’s countermeasures: summarized reasoning, preserved thinking in Fable 5.1, identity checks for unsupported regions.
§ 01The two trends Anthropic leads with
The cyber section opens with two claims and then spends its case studies proving them. The first is about who can attack: “AI has collapsed the labor and tooling gap that used to separate well-resourced, state-sponsored operations from individual operators”, so that “sophistication has stopped being a reliable signal of who is behind an operation”. The second is about how: “A majority of the operations described in this report were enabled by AI via direct execution or orchestration”, meaning multi-agent frameworks running reconnaissance, exploitation and exfiltration while humans picked targets and reviewed what came back.
| Trend | Anthropic’s words | The case it points to |
|---|---|---|
| Sophisticated attacks no longer require sophisticated attackers | “The main distinguishing feature between these classes of actors is no longer sophistication but intent” | A French-speaking hacktivist on stolen API keys (GTG-50029), a credential-harvesting crew (GTG-50014) and a state espionage operator (GTG-20006) ran the same kind of multi-victim campaign |
| AI’s role has become increasingly autonomous | The November 2025 operating model “has now proliferated across every class of actors we investigated” | GTG-20006 ran monitoring agents that rebuilt and redeployed its malware whenever a security product detected it |
| The tooling is public | Offensive agent frameworks such as PentAGI “reproduce much of the same scaffolding for anyone who downloads them” | GTG-50020 and GTG-50029 ran on public frameworks or derivatives |
Anthropic’s own conclusion from this: “The capabilities described in this report should be assumed to be available to any actors who are motivated to use them.”
§ 02The cyber cases
| Group | Who, per Anthropic | What Claude did | The number that sticks |
|---|---|---|---|
| GTG-20006 | Russian espionage, attribution consistent with public reporting on Midnight Blizzard, an operator using the handle JackPoterz | Managed a custom toolkit of implants, a mobile exploitation kit, a credential stealer and a phishing platform through AI-assisted workflows; monitored detections and re-tooled | Targets in Ukrainian and European governments, diplomatic and defense organizations |
| GTG-50014 | ShinyHunters-associated affiliates, one French-speaking operator under three aliases | Ran a credential-harvesting pipeline across a fleet of 10 AWS EC2 workers, decompiling apps and scanning for hardcoded secrets | 1.8 million distinct Android APKs downloaded and scanned |
| GTG-10007 | Chinese-speaking operators likely in Changsha, two of them undergraduates at a Hunan university | The engineering and orchestration layer of an espionage program: exploit foundries, reconnaissance loops, collection fleets | Vulnerability research running agentically around the clock |
| GTG-50020 | Russian-speaking, financially motivated, previously hotel-booking and fintech intrusions | Prompt-injected an AI vendor’s evaluation sandbox to obtain production API keys, then reused them | About thirty AI companies attacked in about four days |
| GTG-50021 | Russian- and Ukrainian-speaking group, one alias kl1zy | Ran a fraudulent reseller of cheap Claude access that proxied traffic elsewhere and installed a credential harvester | Access that was “neither cheap nor actually Claude” |
| GTG-50029 | A single French-speaking hacktivist | Used stolen API keys and a public agent framework against European political parties, media, think tanks and SaaS providers | One person, multi-victim campaign |
§ 03The AI supply chain is now the target
This is the section we would put in front of anyone who runs agents with real credentials. Anthropic’s sentence: “Access to AI in the form of compromised API keys, session tokens, and devices has increasingly become the sole objective of multiple criminal groups.” The stolen access flows to brokers, then to fraudulent reseller networks that rotate through keys until each one is exhausted, and back to attackers who run their operations on other people’s quotas.
The GTG-50020 case is the sharpest. “By injecting malicious instructions into an AI vendor’s automated evaluation sandbox, the actor caused the sandbox to hand over the credentials it held”, including production API keys from several providers. With those keys in hand the actor “automatically switched to using the victim’s keys instead of their own”, then ran a follow-on campaign that “attacked roughly thirty AI companies in about four days with similar techniques”, one working attack path repeated against all thirty. Anthropic calls it “the clearest demonstration to date that the AI supply chain has become a deliberate criminal target”, and is precise about the boundary: “The actor never compromised Anthropic’s own systems”, and the goal of reaching a pre-release model was never realized.
| Pathway | What it looks like | Who it hits |
|---|---|---|
| Spoofed clients | Sites offering discounted frontier access deliver “malicious client side applications often spoofing as popular AI harnesses including Claude Code” that harvest every credential and session token on the device, and keep harvesting after a reset | Developers looking for a cheaper Claude |
| Fraudulent resellers | GTG-50021 sold Claude access that proxied traffic to a different model while its tooling stole the buyer’s Anthropic credentials and resold them | Buyers in unsupported regions and bargain hunters |
| Prompt injection against wrappers | Multiple actors compromised AI wrapper services’ LiteLLM deployments: “they used prompt injection to exfiltrate the production API keys used in their cloud-hosted container environments” | Any service that puts live keys where an agent can be talked into reading them |
| Exposed keys | Legitimate customers who left API keys and session tokens in products, applications and public repositories; the ShinyHunters affiliates industrialized the search across 1.8 million apps | Everyone |
§ 04Distillation: seven labs, one at 151 million exchanges
Anthropic defines the term before naming names: “We define illicit distillation as an industrial-scale, covert campaign to extract a model’s capabilities and replicate them in another model without authorization”, enabled by fake accounts, stolen cards and proxy networks that route around geographic restrictions. It then attributes campaigns to seven China-based labs, all against generally available models; none targeted Mythos 5 or Mythos Preview.
| Lab | Anthropic’s finding | Attributed volume |
|---|---|---|
| Alibaba (Qwen, GTG-16005) | Chain-of-thought extraction to train Qwen 3.5, 3.6 and 3.7; also used Claude to build RL environments; nearly 5,000 fraudulent accounts in the first pool, more than 3,500 behind the peak | Over 151 million exchanges, May to July 2026; peak nearly 3 million per day |
| Moonshot (Kimi, GTG-16002) | Silently forwarded customer requests to Claude and showed users Claude’s answers as Kimi’s; extracted reasoning traces through a cross-session replay of the thinking signature; 5,380 fraudulent accounts | Over 23 million exchanges, May to July 2026 |
| DeepSeek (GTG-16001) | Same replay attack; relayed requests from users on Claude Code, the Claude Agent SDK and OpenCode to Claude Opus without telling them | Over 12.1 million exchanges in 14 days of July 2026 |
| Zhipu (Z.ai, GTG-16006) | Replayed captured reasoning through Claude to clean it for GLM training; 273 fraudulent accounts against Opus 4.8 | 770,609 exchanges through the cleaner in ten June days; over 3 million attributed |
| Xiaomi (MiMo, GTG-16008) | Replayed its own users’ MiMo sessions, often from OpenClaw and OpenCode, through Claude for SFT and RL data; more than 1,500 accounts | More than 400,000 requests |
| SenseTime (GTG-16012) | Bought transcripts of user exchanges with Claude from third-party data vendors; used Claude to write its distillation pipeline | Not stated |
| MiniMax (GTG-16003) | Built a proxy service through a shell company that offers only Anthropic and OpenAI models, none of its own | Not stated |
The relay cases are the ones that will surprise users of those products. Moonshot “silently forwarded customer requests to Claude, instead of processing them using Kimi”, almost 300,000 of them in one ten-day window, most routed to Opus. DeepSeek checked inbound requests for strings that mark coding harnesses and relayed tagged users to Claude Opus; Anthropic’s examples include an employee analyzing internal documentation and an engineer building a case-management system for a municipal Public Security Bureau. The countermeasures Anthropic lists: “Claude now summarizes its internal reasoning before responding, which makes stolen transcripts less useful for training another model”, Fable 5.1’s preserved thinking stops new API accounts from editing the context that precedes a reasoning block, and accounts operating from unsupported countries or reselling access can be required to verify their identity.
§ 05The other five harm areas
| Harm area | Cases | What stands out |
|---|---|---|
| Influence operations | Nine, “originated in Russia, Iran, Turkey, and across the Gulf, South Asia, Africa and Europe, and targeted audiences on six continents” | Influence sold as a service by working ad agencies; “AI as a newsdesk”, where Claude slotted into an existing human-edited pipeline as a sub-editor; several campaigns timed to elections in Moldova and Kenya |
| Surveillance | Ten sections, from a commercial vendor profiling Gulf social-media users to Chinese religious-affairs monitoring and two Iranian units | Lakana 360, a platform for Mali’s state intelligence service built by a single subscriber to monitor “roughly 25 million SIM cards on all three of the country’s national mobile operators” |
| Conventional weapons | Six: a Yemen-based guided-weapons cell, a Russian autonomous drone swarm, Chinese fire-control, electronic-warfare and directed-energy work, and Russian dual-use procurement | Claude drafting fire-control specifications and export-diversion paperwork, not just answering questions |
| Biological misuse | Five case studies, institutions and countries withheld, dual-use acknowledged | Anthropic says it launched Claude Fable 5 with stronger safeguards restricting dual-use biological queries |
| Scams and fraud | One network, GTG-15001: a China-based studio running more than 20 dating apps | 4,700 AI personas engaged at least 25,000 people in two April weeks, about 2.36 million messages, mixed with gig workers at a “3-to-1 AI-to-human ratio” for video calls |
§ 06Which models, and what Anthropic changed
The report is careful about model generations. In the summary: “None of the misuse cases involved the use of Claude Fable or Mythos-class models, with the exception of one illicit distillation case.” In the cyber section: no malicious activity was found on Fable or Mythos, which carry safeguards that Anthropic says greatly reduce their ability to perform harmful cyber tasks. Every case in the report ran on Haiku, Sonnet or Opus.
What changed on Anthropic’s side, per the report: distillation classifiers strengthened with the Fable 5 launch, summarized reasoning in responses, preserved thinking in Fable 5.1, identity verification for accounts that look like resale or unsupported-region access, and in every case the accounts banned, the safeguards updated, and intelligence shared with authorities and industry partners.
§ 07Anthropic’s threat reports so far
| Date | Report | Focus |
|---|---|---|
| March 2025 | Detecting and countering malicious uses of Claude | First public case studies |
| August 2025 | Detecting and countering misuse of AI | Cybercrime, extortion, DPRK employment fraud |
| November 2025 | Disrupting the first reported AI-orchestrated cyber espionage campaign | One state-nexus operation run largely by an agent |
| February 2026 | Detecting and preventing distillation attacks | DeepSeek, Moonshot and MiniMax named |
| September 10, 2026 | Detecting and countering misuse of AI: September 2026 | Seven harm areas, eight months, roughly 40 groups |
§ 08What this means if you run agents on real systems
Our conflict, declared: we build CellCog, an AI employee platform, and our employees run on Claude. Agent, Agent Creative and Agent Team route to Claude Fable 5.1 at the Core and Max tiers, with Gemini 3.8 Flash at Flash, and we publish which model runs where. Anthropic reports no misuse on Fable-class models except the one distillation case, and the cases in this report ran on older generations. That is the reassuring half.
The half that applies to every agent platform, ours included, is the supply chain. Read the four pathways again: a sandbox that held live production keys and could be talked into handing them over; wrapper services whose containers carried the keys an injected prompt could exfiltrate; installers that harvested every token on a developer’s machine; keys left in apps and repositories. None of those is a model failure. Each is a place where a credential sat within reach of something that could be manipulated.
That is the shape we built against. Every CellCog employee runs in its own isolated workspace. Credentials are scoped per employee and enforced on the server, so a prompt that reaches an agent does not reach a key. Every command that reaches your world, a terminal on your machine, your real browser, a connected tool, is classified by the agent before it runs, the platform rejects any command that arrives unclassified, and the consequential ones wait for your approval. The alignment assessment showed why the boundary has to be a separate layer from the model’s judgment; this report shows why it also has to be a separate layer from the model’s credentials.
§ 09What we are watching for
- The report PDF. Anthropic links a PDF edition from the page; if it carries indicators or figures beyond the web version, we fold them in here.
- Responses from the named labs. Alibaba, Moonshot, DeepSeek, Zhipu, Xiaomi, SenseTime and MiniMax have not, as of publication, responded on the record in anything we could read from a primary source.
- The sandbox vendor. Anthropic does not name the AI vendor whose evaluation sandbox was injected; a disclosure from that company would be the first outside account of the case.
- Whether the next report names a Fable-class misuse case. The one-case exception is the number to watch across generations.
§ 10Sources
Anthropic, Detecting and countering misuse of AI: September 2026, September 10, 2026 (read in full the day of publication; every quotation above is from this page, and the PDF edition is linked from it). Anthropic’s Threat Intelligence index lists the earlier reports, including Detecting and preventing distillation attacks (February 2026). Our routing statements reflect CellCog’s configuration as of September 10, 2026.
Q1What is a GTG?
A Generative Threat Group, Anthropic’s internal designator for an actor observed abusing AI. The report uses roughly 40 of them, from GTG-20006 (an actor whose tradecraft Anthropic links to public reporting on Midnight Blizzard) to GTG-16005 (the Alibaba distillation campaign).
Q2What exactly did the sandbox attack do?
Anthropic describes GTG-50020, a Russian-speaking, financially motivated actor with a history of hotel-booking and fintech intrusions. It injected malicious instructions into an AI vendor’s automated evaluation sandbox, which handed over the credentials it held, including production API keys from multiple providers. The actor then used the victim’s keys for further attacks and ran a follow-on campaign against roughly thirty AI companies in about four days, seeking a pre-release Claude model. Every path failed, and the keys were customers’ keys stolen from customers’ environments, not Anthropic’s.
Q3What is illicit distillation, and how big is it?
Anthropic’s definition: ‘an industrial-scale, covert campaign to extract a model’s capabilities and replicate them in another model without authorization’, enabled by fraudulent accounts, stolen cards and proxy networks. Attributed volumes: Alibaba over 151 million exchanges (May to July 2026), Moonshot over 23 million, DeepSeek over 12.1 million in 14 days of July, Zhipu over 3 million, Xiaomi more than 400,000 requests.
Q4Did users of Kimi or DeepSeek know they were talking to Claude?
Anthropic says no. Moonshot ‘silently forwarded customer requests to Claude, instead of processing them using Kimi’, about 300,000 requests over one ten-day window, and DeepSeek relayed requests from users on coding harnesses like Claude Code, the Claude Agent SDK and OpenCode to Claude Opus, in one case an engineer’s specification for a municipal Public Security Bureau case-management system.
Q5Does CellCog run on the models in this report?
Our employees run on Claude. We route Agent, Agent Creative and Agent Team to Claude Fable 5.1 at the Core and Max tiers and Gemini 3.8 Flash at Flash, and we say which model runs where. Anthropic reports no misuse on Fable-class models except one distillation case; the cases here ran on Haiku, Sonnet and Opus. The part of the report that applies to any agent platform is the supply chain: keys and sessions stolen from customer environments, which is why ours are enforced on the server and never handed to the client.




