AI Safety & Security

A 16,000-Request Campaign Targeted OpenAI’s Model Reasoning

OpenAI model targeted by a 16,000-request campaign probing its AI reasoning
A 16,000-request campaign targeted an OpenAI model’s reasoning, illustrating how AI systems can face large-scale, systematic probing.

OpenAI model distillation defences disrupted what the company describes as a coordinated campaign that tried to extract protected reasoning from its models, including a spike of 16,000 extraction-pattern requests across two days in July. The company says the operators did not break its encryption, compromise a database or gain direct access to stored user conversations.

The incident shows a different kind of AI security problem. Instead of stealing model files from a server, attackers can repeatedly interact with a model and try to make it reveal useful reasoning that can be collected as training data for another system.

OpenAI model distillation targeted behaviour, not a database

In its 30 September security report, OpenAI describes the activity as adversarial distillation. Distillation itself is a legitimate technique in which a smaller or secondary model learns from the outputs of a more capable model. The adversarial version uses that process without permission, often at scale and in ways that violate the provider’s terms.

OpenAI says the campaign was designed to extract protected reasoning, meaning internal reasoning traces that the company does not normally expose in the final answer. According to the report, some operators attempted unusual techniques such as moving encrypted reasoning material between conversations and asking another model interaction to decrypt or transcribe it.

The company is explicit that this was not a conventional breach of stored user data. That distinction matters in the wake of broader concern around OpenAI and model cybersecurity risks. A successful model-extraction attempt can be serious without implying that customer conversations were stolen from a database.

OpenAI says activity surged on 24 and 25 July

OpenAI says the earliest activity it observed began on 1 July at relatively low volume. It then recorded high-volume spikes on 24 and 25 July consisting of 16,000 requests using a relevant extraction pattern from more than 4,000 users.

Further investigation, according to the company, identified related prompt-pattern activity across a cluster of more than 15,000 users. OpenAI says it fully disrupted that cluster by 28 July.

The company’s footnote is important: the figures describe extraction attempts, not necessarily successful extraction. A request that matches an attack pattern does not prove the operator obtained useful protected reasoning every time.

OpenAI also says it is unclear whether every operator observed in the period belonged to one actor. It attributes a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi. That attribution is OpenAI’s assessment and has not been independently established by the evidence available in the company’s public report.

Why reasoning traces are valuable training material

A frontier model’s final answer is only one piece of information. The sequence of intermediate reasoning used to reach that answer can show how the system decomposes a problem, evaluates options and corrects itself.

If another developer can gather enough high-quality examples, those traces may become useful synthetic training data. The goal is not necessarily to copy every parameter of the original model. It is to transfer some of the behaviour and capability into another system.

The Frontier Model Forum’s February issue brief on adversarial distillation describes several techniques, including attempts to elicit chain-of-thought reasoning, use models to critique candidate reasoning and generate synthetic training data at scale.

The Forum also draws a line between authorised distillation, which can make smaller and cheaper models, and adversarial distillation that violates access rules or seeks to replicate protected capabilities without the associated safeguards.

The safety concern is that capability can travel without the guardrails

OpenAI argues that extracted reasoning could help train another model without preserving the safety measures surrounding the original system. That is the core security concern.

A frontier lab can spend heavily on monitoring, refusal behaviour, access controls and abuse detection. If another actor learns from the model’s outputs and deploys the resulting capability in a system with weaker safeguards, the capability may spread faster than the controls.

This does not mean distillation automatically produces an equivalent model. Training quality depends on the data collected, the student model, compute and many other choices. The concern is that model interactions themselves can become a route for transferring expensive capability.

Traditional AI-agent data theft risks focus on confidential information. Adversarial distillation is related but distinct: the target is the model’s behaviour and reasoning rather than a conventional store of business data.

The incident shows why rate limits alone are not enough

A large extraction campaign may look like ordinary model usage spread across many accounts. Blocking one user does little if the operator can rotate identities, vary prompts and distribute traffic.

OpenAI says the activity evolved over time and that its response included technical mitigations, improved detection and enforcement against coordinated campaigns. It also shared information with industry partners through the Frontier Model Forum.

This kind of defence increasingly resembles fraud detection. Providers need to identify patterns across accounts, recognise unusual sequences of prompts and distinguish legitimate high-volume research from activity designed to reproduce the model.

That distinction is difficult because many benign users also ask models to explain reasoning, generate synthetic examples or critique solutions. A detector that is too aggressive can block legitimate work, while a detector that is too permissive leaves the extraction route open.

Industry sharing matters because the attack can move between models

Adversarial distillation is not unique to one provider. An operator can test techniques against several frontier systems and adapt when one company changes its defences.

That gives competing AI labs an unusual reason to cooperate. They may disagree on products and policy while still sharing threat indicators about coordinated extraction campaigns.

The Frontier Model Forum already operates an information-sharing mechanism for vulnerabilities, threats and capabilities of concern. OpenAI’s decision to share details of this campaign fits that model: one provider’s incident can become an early warning for others.

The public report leaves important questions unanswered

OpenAI has not published everything it knows about the campaign, and doing so could make future attacks easier. The public report therefore cannot independently establish the full attribution, the amount of useful reasoning obtained or the capability gained by any secondary model.

Those limits should remain visible. “Disrupted” means OpenAI says it stopped the observed cluster, not that adversarial distillation as a technique has been solved.

The company itself expects attempts to become more sophisticated as frontier models improve. That is plausible because the economic incentive grows with model capability. If a competitor can cheaply learn from a system that cost far more to train, extraction becomes attractive even when only part of the behaviour transfers.

AI models are becoming targets in their own right

Cybersecurity usually protects data, money and access. Frontier AI adds another asset: capability.

A model’s useful behaviour can be queried through an API, which means the interface that creates the product’s value can also become an extraction surface. Providers have to serve legitimate users while detecting when those same interactions are being assembled into a training pipeline.

OpenAI’s July campaign is therefore less like somebody breaking into a vault and more like somebody learning how to drain value through the service window.

The 16,000 requests over two days make the scale visible, but the lasting issue is structural. As frontier systems become more capable, keeping the model weights secure will not be enough. Labs will also need to defend the behaviour that users are allowed to access one request at a time.

About the Author

Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.