OpenAI has published an account of a campaign aimed at extracting its models' protected reasoning. In its Sept. 30 report, the company says it disrupted the related activity by July 28, following attempts that began July 1.
Protected reasoning is a model's internal record of working through a task. OpenAI describes the campaign as adversarial distillation: unauthorized use of a model's outputs or reasoning to help reproduce or improve another model.
What the extraction attempts involved
OpenAI says operators copied encrypted reasoning from one conversation and asked a model in another conversation to decrypt and transcribe it. The company says the operators did not break its encryption, compromise a database or gain direct access to stored user conversations.
The report describes spikes on July 24 and 25 totaling 16,000 requests from more than 4,000 users. A wider investigation identified related prompt patterns across more than 15,000 users. OpenAI specifies that these figures count attempted extractions, rather than necessarily successful ones.
A limited attribution and continuing defenses
OpenAI attributes a core group of the activity to individuals associated with Moonshot AI, the developer of Kimi. It says it is unclear whether all observed operators came from a single actor. That attribution does not establish that every attempt belonged to Moonshot AI.
The company says it banned or restricted fraudulent accounts, strengthened signup controls and added protections for hidden reasoning. It also reports closing a pathway that let someone possessing another user's encrypted reasoning replay it and recover its contents, and adding checks for streamed output that might expose reasoning.
OpenAI says it worked with third-party providers to disrupt related accounts and shared findings through the Frontier Model Forum. Its report says mitigation and investigation continue, including work to extend protections to partner-hosted deployments.