AI · 4h ago
OpenAI said it disrupted a coordinated distillation campaign that began July 1, spiked on July 24-25, and involved more than 15,000 users; it attributed a core cluster of the activity to people associated with Moonshot AI. The company said the campaign was aimed at extracting protected reasoning from its models, not at breaking into stored conversations or a database.
OpenAI says the operators worked by scaling ordinary model interactions until hidden reasoning reappeared in a form the requester could see. It also closed a replay path for some encrypted reasoning and added checks to hold streamed output that might expose that internal trace, which matters because the asset under attack here is model behavior itself, not records at rest.
For teams running chatbots, copilots, or agents that stream output or expose reasoning traces, the exposure sits at the interaction layer: if an attacker can query at scale, they may be able to harvest model behavior even when your back end stays untouched. The reporting does not settle how broadly similar protections generalize across vendors, but it does show that normal incident response aimed at data theft misses this class of loss.
3 sources covering this story
OpenAI Disrupts Reasoning Extraction Campaign Linked to Moonshot AI Associates
OpenAI disrupted a coordinated campaign attributed to Moonshot AI associates that sought to extract protected reasoning at scale.
Srsly Risky Biz: “Rogue AI” isn’t going anywhere
Amberleigh Jack and James Wilson chat about OpenAI agents’ recent escapades into Australian government websites.
Admits its agents side-swiped four Australian government sites
Part of the PlainSec briefing for 2026-10-01