Skip to content

Security1 publisher2 min readPublished

OpenAI says a distillation campaign got its models to decrypt their own protected reasoning

OpenAI says 15,000 suspicious users had its models decode encrypted reasoning copied from other chats before it cut them off on July 28. The encryption held throughout, and the only barrier left was what the model would agree to decode.

The Watch · Security desk

Illustration accompanying OpenAI says a distillation campaign got its models to decrypt their own protected reasoning

What happened

  • Low-level activity began July 1 and peaked on July 24 and 25, when OpenAI counted 16,000 prompts from 4,000 users matching one extraction pattern.
  • OpenAI ties a "core cluster" of the activity to people working for Moonshot AI, while saying it cannot tell whether all of it is related.
  • Beyond banning the accounts, OpenAI says it tightened signup and infrastructure controls and expanded network monitoring.
  • Outside researchers reported a similar vulnerability to OpenAI in August, after the July 28 disruption.

Compiled by The WatchSomething wrong?How this is made

Why it matters

  • precedent Any lab that returns encrypted reasoning to users will now be expected to show that a different conversation cannot get the model to decode it.
  • constraint At four prompts per user, no single account stood out, so detection has to match the shared pattern across thousands of accounts.
  • exposure On OpenAI's account, ChatGPT customers are not the exposed party, because no database or stored user conversation was reached and the material taken was the model's own reasoning.
  • constraint Outside the groups OpenAI has briefed, the Moonshot attribution has to be taken on trust until OpenAI or Moonshot puts evidence on the record.

OpenAI's account is that the cryptography held. "The operators did not break our encryption, compromise a database, or gain direct access to stored user conversations," the company wrote in an unsigned blog post [6]. "Instead, they manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester in a coordinated, scaled manner that violated our terms of service." [7] The operators copied encrypted reasoning out of one conversation, then asked the model in a separate conversation to decrypt it and write it out in plain text [4]. OpenAI called the method "novel" [5].

The fix shows where the check was missing. OpenAI says it closed a bug that let users take encrypted data from one conversation and decrypt it in another [16]. Until then, the reasoning stayed encrypted wherever the operators carried it, and a second session would still turn it back into text when asked [4]. Encrypting the output protected it here only as far as the service refused to decrypt it for a conversation that did not produce it [4][16]. OpenAI says the same vulnerability exists in other AI models, and it has shared details of the incident with groups including the Frontier Model Forum [14][15].

The traffic was spread thin. At the July 24-25 peak, matching prompts averaged four per user [1], so no single account had to look heavy. Suspicious users then grew 3.75-fold in three to four days [2]. First sighting to full disruption took 27 days [3]. That spread fits the pattern Google and other security firms describe in Chinese distillation work: thousands of individual accounts bought on black or gray markets for models like ChatGPT and Claude, then used to flood those models with millions of prompts [11].

CyberScoop reported that the post cites no technical evidence or reasoning for naming Moonshot AI [9]. OpenAI told the outlet it was withholding further detail "for security reasons" [12]. CyberScoop said it had asked Moonshot for comment [13]. The technique and the counts come from OpenAI's post; the Moonshot link is OpenAI's assertion. American AI companies and the U.S. government have previously accused Chinese firms, Moonshot among them, of "systematic" distillation attacks on their latest models [10].

What to watch

  • Whether other labs confirm the cross-conversation decryption flaw in their own models, and say how they closed it.
  • Technical indicators from OpenAI, or a response from Moonshot AI, that would let outsiders test the attribution.
  • Details of the outside researchers' August report, including whether their similar flaw still worked after OpenAI's changes.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories