Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentTutorialsPricingBlogContact

OpenAI Distillation Campaign: What It Says About Moonshot

At a glanceQuick answers
What did OpenAI disclose?
A campaign from July 1 to July 28, 2026 that tried to extract its models’ hidden reasoning at scale, now disrupted and shared with the Frontier Model Forum.
Who does OpenAI blame?
A core cluster tied to individuals associated with Moonshot AI, the maker of Kimi.
What is the one caveat to carry?
The numbers are attempts, and OpenAI does not say how much reasoning, if any, was actually recovered.
Editorial data illustration on a near-white ground titled OpenAI disrupts a distillation campaign: a navy padlocked box labeled reasoning with a dotted teal stream leaking out through an amber-marked tube, and two large numbers, 16,000 extraction requests on July 24-25 and 15,000+ users in the cluster
Fig 0Locked reasoning, leaking through a straw. Made by CellCog's image agent, running GPT Image 2.5.

OpenAI says a coordinated campaign spent July trying to pull the hidden reasoning out of its models, and it points at Moonshot AI. In a post published September 30, 2026, OpenAI writes: “we attribute a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi.” It adds that “It is unclear whether all operators we observed during the relevant time period originated from a single actor.”

This page is read from OpenAI’s post and the independent research paper it links. It is the second distillation allegation against Moonshot this month; Anthropic’s came on September 10, in the threat report we covered.

On this page · 8 sectionsOpen
  1. What OpenAI says happened
  2. How the extraction worked
  3. What OpenAI changed
  4. What the post does not tell you
  5. What this means if you run agents on real work
  6. What we are watching for
  7. Update log
  8. Sources
Key points6 · 5 min full read
  1. On September 30, 2026, OpenAI said it identified and disrupted a coordinated campaign to extract protected reasoning from its models, which it calls adversarial distillation.
  2. OpenAI attributes a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi, while saying it is unclear whether all operators came from one actor.
  3. Activity began July 1, spiked on July 24 and 25 with 16,000 extraction requests from more than 4,000 users, and a related cluster of more than 15,000 users was fully disrupted by July 28.
  4. The technique copied encrypted reasoning from one conversation and asked a model in another conversation to decrypt and transcribe it; no encryption was broken and no database was compromised.
  5. OpenAI credits independent researchers whose August paper showed encrypted reasoning blocks are interchangeable across sessions, users and models at Anthropic, OpenAI and Google.
  6. The counts are attempted extractions, not necessarily successful ones, per OpenAI’s own footnote.

§ 01What OpenAI says happened

Date Event
July 1 Activity begins at low volume
July 24 to 25 Spikes: 16,000 requests with an extraction pattern from more than 4,000 users
By July 28 Related activity across a cluster of more than 15,000 users fully disrupted
August 10 Independent researchers publish the cross-model reasoning attack
September 30 OpenAI publishes its account
Table 1The campaign, dated, per OpenAI’s September 30, 2026 post
From first request to disclosure, July 1 to September 30, 2026Timeline of the distillation campaign: July 1 start, July 24 spike highlighted, July 28 disrupted, August 10 paper, September 30 disclosureJul 1Activity beginsJul 2416,000-request spikeJul 28Cluster disruptedAug 10Research paperSep 30OpenAI disclosesFrom first request to disclosure, July 1 to September 30, 2026Timeline of the distillation campaign: July 1 start, July 24 spike highlighted, July 28 disrupted, August 10 paper, September 30 disclosureJul 1Activity beginsJul 2416,000-request spikeJul 28Cluster disruptedAug 10Research paperSep 30OpenAI discloses
Fig 1From first request to disclosure, July 1 to September 30, 2026

The numbers come with OpenAI’s own footnote: “These figures describe attempted, not necessarily successful, extractions.” The post does not say how much reasoning, if any, was recovered.

§ 02How the extraction worked

Nobody broke in. “The operators did not break our encryption, compromise a database, or gain direct access to stored user conversations.” Instead, OpenAI says, they copied encrypted reasoning from one conversation and asked a model in another conversation to decrypt and transcribe the hidden content.

That is the attack class the linked paper describes. Stealing Reasoning Traces from Proprietary LLM APIs, submitted August 10 by Alexander Panfilov, David Schmotz and six co-authors, found that encrypted reasoning blocks are “fully compatible and interchangeable across different sessions, users, and models” inside one provider’s ecosystem. Injecting a strong model’s encrypted trace into a weaker, less safeguarded model from the same provider made the weaker one print it in plain text. The authors demonstrated it against Anthropic, OpenAI and Google.

§ 03What OpenAI changed

  • Banned or restricted fraudulent accounts and tightened signup and infrastructure controls.
  • Closed the pathway that let someone holding another user’s encrypted reasoning replay it and recover its contents.
  • Added checks that detect and hold streamed output that might expose reasoning.
  • Worked with third-party services the activity moved through, and shared findings through the Frontier Model Forum and government channels.

Its warning to everyone else is one sentence: “Systems that support portable or replayable reasoning artifacts may face related risks.”

§ 04What the post does not tell you

  • How much got out. The counts are attempts; there is no figure for successful extraction.
  • Who exactly. The attribution is to “individuals associated with” Moonshot, not to the company as an institution, and not to every operator.
  • Moonshot’s side. We found no public response from Moonshot AI as of 17:45 UTC on September 30.
  • The evidence. OpenAI does not publish the signals behind the attribution.

§ 05What this means if you run agents on real work

The reasoning a model hides is now a thing worth stealing, and the theft here ran through ordinary product features: carrying state from one conversation to the next, and handing it to a cheaper model. Any platform that passes model state between sessions, users or models is carrying the same kind of artifact. The useful questions are about the layer around the model: whether one worker’s context can reach another’s, and whether anything that leaves the workspace passes a gate.

§ 06What we are watching for

  • A response from Moonshot AI or Kimi’s official accounts.
  • Whether Anthropic or Google publish matching findings through the Frontier Model Forum.
  • Any change to how providers return encrypted reasoning to API clients.

§ 07Update log

  • September 30, 2026: page opened, from OpenAI’s post of the same day and the August 10 paper.

§ 08Sources

Frequently asked6 questions

Q1What is adversarial distillation?

OpenAI defines it as the systematic and unauthorized use of one model’s outputs or reasoning to help train, reproduce or improve another model. Hidden reasoning is valuable because it can reveal what the final answer leaves out.

Q2How did the extraction work?

Operators copied encrypted reasoning from one conversation and asked a model in another conversation to decrypt and transcribe it. OpenAI says they did not break its encryption, compromise a database or reach stored user conversations.

Q3Did Moonshot AI respond?

We found no public response from Moonshot AI to OpenAI’s post as of 17:45 UTC on September 30, 2026. Anthropic made a separate distillation allegation against Moonshot on September 10.

Q4What did OpenAI change?

It banned or restricted accounts, tightened signup controls, closed a path that let someone replay another user’s encrypted reasoning, added checks that hold streamed output which might expose reasoning, and worked with third-party services the activity moved through.

Q5Does this affect other AI providers?

OpenAI says the manipulation is not unique to its models and that systems with portable or replayable reasoning artifacts may face related risks. The researchers’ paper demonstrated the underlying attack at Anthropic, OpenAI and Google.

Q6Where does CellCog fit?

We do not train models and we do not route work to Kimi; our AI employees run on Claude Opus 5.5. We cover this because the security of the models underneath every agent platform, ours included, is the layer our customers inherit.

Published 30 September 2026 All Trust, permissions & security →