OpenAI Accuses Moonshot AI of Stealing Model Reasoning
OpenAI says 4,000 users sent 16,000 prompts to extract its models' encrypted reasoning, and links part of the campaign to people tied to Moonshot AI.

OpenAI says it shut down a campaign that pulled the hidden reasoning out of its models by prompting, not by breaking anything — and it has linked part of that activity to people associated with Moonshot AI, the Chinese startup behind Kimi. The company calls the technique "adversarial distillation."
What OpenAI says happened
According to the report, the campaign began on July 1 at low volume and then spiked: on July 24 and 25 alone, more than 4,000 users fired some 16,000 requests using one particular prompting technique. Further investigation found related prompt patterns across a cluster of more than 15,000 users. OpenAI says it fully shut the operation down by July 28.
The method is the part worth reading twice. Operators copied the encrypted reasoning output from one conversation, pasted it into a separate conversation, and asked a different model instance to decrypt and transcribe it. OpenAI encrypts the chain-of-thought its reasoning models produce precisely so competitors cannot read it.
OpenAI also stresses what it says did not happen: the operators never broke its encryption, never accessed stored user conversations, and never breached a database. They manipulated ordinary model interactions with designed prompts to make hidden reasoning visible.
Why the label matters
Distillation itself is routine: a smaller student model is trained on a larger teacher's outputs to inherit its capabilities cheaply. The report frames "adversarial" distillation as the systematic, unauthorized use of one model's outputs or reasoning to train or reproduce a rival, in violation of terms of service.
OpenAI's stated reason for caring is that protected reasoning is "the model's internal record for working through a task," so extracting it can reveal information withheld from the final answer and help others reproduce advanced capabilities without the original safety guardrails. Caroline Zier, who leads OpenAI's strategic national-security policy work, is quoted drawing the line narrowly: "Our concern is about violation of our terms of service, not open models or legitimate distillation."
The company says it banned the offending accounts, tightened sign-up verification, added protections around hidden reasoning, and shared findings through the Frontier Model Forum and government information-sharing channels. Moonshot AI has not publicly responded to the allegations, and the accusations have not been independently verified.
What this means if you build on these APIs
Two practical consequences. First, sign-up verification is getting tighter on the back of this, which is the kind of change that lands on legitimate developers as friction before it lands on anyone else. Second, this is an accusation from one party with a direct commercial interest, so treat it as a disclosed investigation rather than an established fact until Moonshot responds or a third party checks the claim.
It also is not isolated. The report notes Anthropic recently accused several Chinese AI developers, including Moonshot and Alibaba, of large-scale distillation attempts against its Claude models, and that Google has made similar complaints. If you are evaluating a cheap model whose capabilities look surprisingly close to a frontier system, provenance is now a diligence question, not a philosophical one.
What to watch
Watch for a Moonshot response, and for whether any lab publishes evidence that a third party can check. The harder question underneath is whether chain-of-thought can be kept secret at all: the current industry answer — encrypt the reasoning, ban the accounts — reads like the opening of an arms race rather than a fix.
More from DangMua