Skip to content

Build3 publishers3 min readPublished

A reasoning-replay attack OpenAI blocked on its own API still worked on Azure in September

OpenAI says it shut down a 15,000-account campaign to extract its models' hidden reasoning by July 28. A September 13 retest still pulled that reasoning verbatim through Azure, so the protection a team gets depends on which platform serves the model.

The Engineer · Build desk

Illustration accompanying A reasoning-replay attack OpenAI blocked on its own API still worked on Azure in September

What happened

  • The attackers copied encrypted reasoning out of one conversation and asked a model in a separate conversation to decrypt it and write it out.
  • Activity started at low volume on July 1, then spiked on July 24 and 25 to 16,000 requests from more than 4,000 users.
  • OpenAI links a core group to people associated with Moonshot AI, the maker of Kimi, but says it is unclear whether all the actors trace back to one source.
  • OpenAI says it banned fraudulent accounts, tightened sign-ups, closed the hole for reading other users' encrypted reasoning, and now holds back streamed output that might reveal it.
  • Anthropic separately reported Moonshot relaying almost 300,000 requests through 5,380 fraudulent accounts in ten days, replaying Claude's reasoning signatures in new sessions.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • contradiction OpenAI describes the replay hole as closed, while the researchers say many fixes are brittle matches on specific request patterns, so a reworded request may pass where the known one is blocked.
  • constraint Closing the ciphertext replay path leaves the notepad-tool route open, so providers need a separate set of tool-output checks before hidden reasoning can be called protected.
  • precedent Any service that hands clients encrypted state to carry back now faces a documented attack class: a shared key lets one tenant's block be read through another session or a cheaper model.

The design under attack is a reasonable one. A reasoning model's internal record goes to the client only as an encrypted block, and the client sends that block back with each request so the provider does not have to store it [7]. The defect was key scope. Joachim Schaeffer's team, in a paper dated August 10 that tested OpenAI, Anthropic and Google, found the blocks were encrypted with shared keys [25][8]. They could move between sessions, between users, and between models from the same provider [8]. That lets a weaker, cheaper model in the same family act as a "decryption oracle" that prints the stronger model's hidden reasoning word for word [8]. OpenAI says the encryption was not broken and no database was compromised [9]. The hidden steps are worth the effort. OpenAI says they can hold information deliberately kept out of the final answer, and can help others recreate a model's capabilities [10]. The traffic was spread thin. The late-July spike works out to at most four requests per user over two days [1], and OpenAI found the wider network through related patterns across accounts [3]. Anthropic's figures come to about 56 requests per account over ten days [2]. OpenAI counts its cases as attempted extractions, not necessarily successful ones [4]. The post does not give a success rate, the models targeted, or how many users were linked to Moonshot [11]. Moonshot denied in July that Kimi K3 was built by distillation, and OpenAI President Greg Brockman said at the time it was "too early" to tell [22]. Beyond the bans and the closed replay path, OpenAI says it strengthened protections for hidden reasoning across users, workspaces, organizations and model families [13]. If that means each block is now bound to the tenant and model that produced it, it is the correct fix for a shared-key problem. OpenAI credited the researchers by name and says their findings sped up its countermeasures [14]. The campaign also crossed hosts. OpenAI says it worked with third-party providers to disrupt accounts whose activity moved through their services [24]. The researchers retested on September 13. The attack was blocked on OpenAI's and Anthropic's own APIs [15]. On Azure it worked against every OpenAI model they tried, including GPT-6 Astra, and against Anthropic models up to Sonnet 5 [15]. One attempt was enough to pull the reasoning out verbatim [15]. That was 47 days after OpenAI says the campaign was fully disrupted [3]. "Same models, but different protections depending on which platform serves them," Schaeffer said [16]. The researchers call the fixes so far piecemeal and superficial, and say some reached cloud platforms only days late [17]. OpenAI's stated next step is to give partner-hosted deployments the same protections as its first-party tools [18]. A second route skips the ciphertext entirely. Developer Can Bölük showed publicly that a model given a virtual notepad tool, and told to write its reasoning there, leaves that reasoning where the user can read it [19]. The researchers say this worked on every OpenAI model and on Opus 4.8 and Sonnet 5. Only Opus 5, Fable 5 and Fable 5.1 kept their reasoning back [20]. The researchers believe the output would likely be as useful for distillation as the decrypted traces [20]. No block is replayed, so a check on replayed blocks has nothing to match. OpenAI says additional checks are needed to prevent tool-output attacks [18]. I think a team serving these models through a cloud host has to treat reasoning protection as a property of that host's deployment, and test it there. The same applies to any system that hands clients encrypted state to carry back: the key has to be scoped to the tenant that produced the block.

What to watch

  • Whether Azure ships the partner-hosted parity OpenAI names as its next step, and whether a fresh retest by Schaeffer's team then fails there.
  • Tool-output checks against the notepad method, which in the researchers' tests only Opus 5, Fable 5 and Fable 5.1 resisted.
  • Retest results for cloud hosts other than Azure that resell the same OpenAI and Anthropic models.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories