OpenAI says a coordinated campaign spent July trying to pull the hidden reasoning out of its models, and it points at Moonshot AI. In a post published September 30, 2026, OpenAI writes: “we attribute a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi.” It adds that “It is unclear whether all operators we observed during the relevant time period originated from a single actor.”
This page is read from OpenAI’s post and the independent research paper it links. It is the second distillation allegation against Moonshot this month; Anthropic’s came on September 10, in the threat report we covered.
On this page · 8 sectionsOpen
- On September 30, 2026, OpenAI said it identified and disrupted a coordinated campaign to extract protected reasoning from its models, which it calls adversarial distillation.
- OpenAI attributes a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi, while saying it is unclear whether all operators came from one actor.
- Activity began July 1, spiked on July 24 and 25 with 16,000 extraction requests from more than 4,000 users, and a related cluster of more than 15,000 users was fully disrupted by July 28.
- The technique copied encrypted reasoning from one conversation and asked a model in another conversation to decrypt and transcribe it; no encryption was broken and no database was compromised.
- OpenAI credits independent researchers whose August paper showed encrypted reasoning blocks are interchangeable across sessions, users and models at Anthropic, OpenAI and Google.
- The counts are attempted extractions, not necessarily successful ones, per OpenAI’s own footnote.
§ 01What OpenAI says happened
| Date | Event |
|---|---|
| July 1 | Activity begins at low volume |
| July 24 to 25 | Spikes: 16,000 requests with an extraction pattern from more than 4,000 users |
| By July 28 | Related activity across a cluster of more than 15,000 users fully disrupted |
| August 10 | Independent researchers publish the cross-model reasoning attack |
| September 30 | OpenAI publishes its account |
The numbers come with OpenAI’s own footnote: “These figures describe attempted, not necessarily successful, extractions.” The post does not say how much reasoning, if any, was recovered.
§ 02How the extraction worked
Nobody broke in. “The operators did not break our encryption, compromise a database, or gain direct access to stored user conversations.” Instead, OpenAI says, they copied encrypted reasoning from one conversation and asked a model in another conversation to decrypt and transcribe the hidden content.
That is the attack class the linked paper describes. Stealing Reasoning Traces from Proprietary LLM APIs, submitted August 10 by Alexander Panfilov, David Schmotz and six co-authors, found that encrypted reasoning blocks are “fully compatible and interchangeable across different sessions, users, and models” inside one provider’s ecosystem. Injecting a strong model’s encrypted trace into a weaker, less safeguarded model from the same provider made the weaker one print it in plain text. The authors demonstrated it against Anthropic, OpenAI and Google.
§ 03What OpenAI changed
- Banned or restricted fraudulent accounts and tightened signup and infrastructure controls.
- Closed the pathway that let someone holding another user’s encrypted reasoning replay it and recover its contents.
- Added checks that detect and hold streamed output that might expose reasoning.
- Worked with third-party services the activity moved through, and shared findings through the Frontier Model Forum and government channels.
Its warning to everyone else is one sentence: “Systems that support portable or replayable reasoning artifacts may face related risks.”
§ 04What the post does not tell you
- How much got out. The counts are attempts; there is no figure for successful extraction.
- Who exactly. The attribution is to “individuals associated with” Moonshot, not to the company as an institution, and not to every operator.
- Moonshot’s side. We found no public response from Moonshot AI as of 17:45 UTC on September 30.
- The evidence. OpenAI does not publish the signals behind the attribution.
§ 05What this means if you run agents on real work
The reasoning a model hides is now a thing worth stealing, and the theft here ran through ordinary product features: carrying state from one conversation to the next, and handing it to a cheaper model. Any platform that passes model state between sessions, users or models is carrying the same kind of artifact. The useful questions are about the layer around the model: whether one worker’s context can reach another’s, and whether anything that leaves the workspace passes a gate.
§ 06What we are watching for
- A response from Moonshot AI or Kimi’s official accounts.
- Whether Anthropic or Google publish matching findings through the Frontier Model Forum.
- Any change to how providers return encrypted reasoning to API clients.
§ 07Update log
- September 30, 2026: page opened, from OpenAI’s post of the same day and the August 10 paper.
§ 08Sources
- OpenAI, Disrupting a coordinated model-distillation campaign, September 30, 2026.
- Panfilov, Schmotz, Shumailov, Beurer-Kellner, Schaeffer, Prabhu, Geiping, Andriushchenko, Stealing Reasoning Traces from Proprietary LLM APIs, arXiv, August 10, 2026.
- Our record of Anthropic’s September threat report.
Q1What is adversarial distillation?
OpenAI defines it as the systematic and unauthorized use of one model’s outputs or reasoning to help train, reproduce or improve another model. Hidden reasoning is valuable because it can reveal what the final answer leaves out.
Q2How did the extraction work?
Operators copied encrypted reasoning from one conversation and asked a model in another conversation to decrypt and transcribe it. OpenAI says they did not break its encryption, compromise a database or reach stored user conversations.
Q3Did Moonshot AI respond?
We found no public response from Moonshot AI to OpenAI’s post as of 17:45 UTC on September 30, 2026. Anthropic made a separate distillation allegation against Moonshot on September 10.
Q4What did OpenAI change?
It banned or restricted accounts, tightened signup controls, closed a path that let someone replay another user’s encrypted reasoning, added checks that hold streamed output which might expose reasoning, and worked with third-party services the activity moved through.
Q5Does this affect other AI providers?
OpenAI says the manipulation is not unique to its models and that systems with portable or replayable reasoning artifacts may face related risks. The researchers’ paper demonstrated the underlying attack at Anthropic, OpenAI and Google.
Q6Where does CellCog fit?
We do not train models and we do not route work to Kimi; our AI employees run on Claude Opus 5.5. We cover this because the security of the models underneath every agent platform, ours included, is the layer our customers inherit.
