OpenAI says a campaign tied to Moonshot AI associates replayed encrypted reasoning to other model paths to decrypt it: 16,000 requests at the July peak.
OpenAI says it found and shut down a coordinated campaign that tried to pull the hidden reasoning out of its models so the output could train someone else's. In a disclosure published on 30 September 2026, the company said a "core cluster" of the activity was tied to individuals associated with Moonshot AI, the Beijing lab behind the Kimi models. Nobody broke OpenAI's encryption. According to OpenAI, the operators copied encrypted reasoning from one conversation and asked a model in another conversation to decrypt it and write it out.
This is OpenAI's allegation. The Hacker News noted that OpenAI did not publish technical evidence for the attribution, and OpenAI itself left open whether all of the activity came from one actor. As of 1 October, reports from The Hacker News, BankInfoSecurity and Parameter said Moonshot AI had not responded to OpenAI's post or to requests for comment.
What OpenAI says happened
The timeline, as OpenAI and the outlets covering it describe it:
- 1 July 2026: activity consistent with adversarial distillation begins at low volume.
- 24–25 July: the campaign spikes to about 16,000 attempted requests from more than 4,000 users.
- 28 July: OpenAI says it disrupted the campaign. Related prompt patterns were eventually found across more than 15,000 users.
- 30 September: OpenAI publishes its account of the campaign. Coverage follows on 1 October.
OpenAI calls the technique adversarial distillation: the systematic, unauthorized use of one model's outputs to train, reproduce or improve another model. Its specific concern is reasoning. Reasoning models think before they answer, and OpenAI does not show that chain of thought to users. Over the API it can hand back an encrypted copy of the reasoning so a developer can pass it into the next turn without ever reading it.
In OpenAI's words, the operators "manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester in a coordinated, scaled manner that violated our terms of service."
The replay path OpenAI describes: the encryption held, but a model path holding the key decrypted reasoning for the wrong requester.
The replay flaw underneath it
The mechanism is not new. In August 2026, researchers from MATS Research and the ELLIS Institute Tübingen, among others, published a study showing that encrypted reasoning traces from Claude, Gemini and GPT were "fully compatible and interchangeable across different sessions, users, and models within a provider's ecosystem." Their attack, as quoted by The Hacker News: "By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly."
ThreatFrontier covered that research when it came out, in LLM Reasoning Block Replay Vulnerability. The point then was that the encryption protected the reasoning in transit but not from the provider's own models, which hold the key. OpenAI's disclosure is the first public account from a provider of that class of attack being run at scale against it.
The encryption was never the weak part. The weak part was that any model holding the key would decrypt an encrypted reasoning blob for whoever asked, and nothing checked that the requester had produced it.
What OpenAI changed
OpenAI and the coverage list these steps:
- banned the accounts involved and strengthened sign-up controls
- closed the pathway that let encrypted reasoning be replayed and decrypted
- added checks that detect and block streamed output that exposes reasoning
- expanded monitoring, and worked with third-party providers to disable linked accounts
- briefed other AI developers through the Frontier Model Forum, and briefed government
OpenAI framed the risk as more than lost revenue: "Extracted reasoning could be used to train another model without preserving the safeguards applied to the original model's user-facing outputs."
The second accusation in a month
Moonshot was already named. On 11 September 2026, Anthropic said seven China-based labs, Moonshot among them, had run distillation campaigns against Claude. It alleged that Moonshot rerouted Kimi users' requests to Claude, showed them Claude's answers and kept the exchanges for training, relaying almost 300,000 customer requests over 10 days through 5,380 fraudulent accounts. The Hacker News reported no response from Moonshot to those claims either.
Two providers have now made similar claims about the same lab within three weeks. Neither has published the evidence behind its attribution, and neither claim has been tested anywhere outside the companies making it.
What this means for teams building on reasoning APIs
The specific pathway is OpenAI's to close, and it says it has. The pattern applies to anyone who builds on reasoning models or ships one:
- Treat encrypted reasoning as sensitive, not as an opaque token. Anything the provider can decrypt, some model path can be persuaded to decrypt. Do not log it, cache it across tenants or pass it between users.
- Bind reasoning to its origin. If you run a model gateway or an agent platform that forwards reasoning items, tie each item to the session and account that produced it, and reject items that arrive from anywhere else.
- Watch for extraction-shaped traffic. A sudden wave of accounts sending near-identical prompts, or requests that ask a model to transcribe or decode material it did not produce, is the signal OpenAI says it used.
- Expect your outputs to be someone's training data. If you expose a model to the public, rate limits and account vetting are distillation controls as well as abuse controls.
Sources
- OpenAI: Disrupting a coordinated model distillation campaign
- The Hacker News: OpenAI Disrupts Reasoning Extraction Campaign Linked to Moonshot AI Associates
- BankInfoSecurity: OpenAI accuses Moonshot AI of coordinated model distillation
- Parameter: OpenAI Accuses Moonshot AI of Attempting to Extract Proprietary Model Reasoning
- The Hacker News: Anthropic Says Seven China-Based AI Labs Ran Industrial-Scale Claude Distillation Attacks