OpenAI said it has identified and disrupted a coordinated campaign aimed at extracting protected reasoning from its artificial intelligence models.
The company linked a core cluster of the activity to individuals associated with Moonshot AI. The Chinese startup is the developer of the Kimi chatbot.
OpenAI described the activity as "adversarial distillation." This is when one AI model's outputs or reasoning are used without permission to help train or improve another model.
How the Campaign Unfolded
The activity started on July 1 at a low volume, according to OpenAI. It then spiked on July 24 and 25, with 16,000 requests coming from more than 4,000 users.
Further investigation found related activity across a cluster of more than 15,000 users. OpenAI said it fully disrupted the campaign by July 28.
Protected reasoning is the model's internal record of how it works through a task. OpenAI said extracting it can reveal information held back from the final answer and help others copy the model's abilities.
The company said operators did not break its encryption, access its databases or reach stored user conversations. Instead, they manipulated interactions with the models so hidden reasoning appeared in a form they could see.
In one method, operators copied encrypted reasoning from one conversation. They then asked a model in a separate conversation to decrypt and transcribe it.
Independent security researchers also reported related weaknesses through responsible disclosure. OpenAI said it confirmed the attack paths they found were real.
OpenAI's Response and Attribution
OpenAI said it was unclear whether all of the operators came from a single actor. It still attributed the core cluster to people tied to Moonshot AI.
The company said it banned or restricted fraudulent accounts and tightened signup and infrastructure controls. It also expanded monitoring for related networks.
OpenAI closed a pathway that let someone holding another user's encrypted reasoning replay it and recover its contents. It also added checks to detect and hold streamed output that might expose reasoning.
When the activity moved through third-party services, OpenAI worked with those providers to disrupt the accounts involved. The company did not name the providers.
OpenAI said it shared its findings with other AI developers through the Frontier Model Forum. It also shared them through government information-sharing channels.
The company said adversarial distillation poses safety and national security risks. It said copied reasoning could be used to train another model without the safeguards built into the original.
The report comes weeks after rival Anthropic accused several Chinese AI developers, including Moonshot AI and Alibaba, of secretly using its Claude model to train their own systems.
OpenAI said its work is not finished and it will keep improving tool defenses, classifiers and model refusals. Moonshot AI did not immediately respond to a request for comment from CNBC.