Build1 distinct publisher3 min readUpdated
One key scheme across a provider's whole model family means the cheapest model in it can be asked to print the expensive one's chain of thought, according to a summary of new research.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Follow any of these and your For You feed starts watching them — no settings page required.
The scope of a key is a design decision, and this one was made against a requirement nobody appears to have written down. The stated need is continuity: the client holds an opaque block and hands it back so a later turn can build on earlier reasoning, with the server keeping the only keys [2]. A key bound to one session, or one paying customer, meets that need in full. What the summary describes instead is a single logical scheme spanning all models, all sessions and all users inside a provider [4], so that a token minted by ChatGPT is readable by GPT-3.5 and a Claude Instant token is readable by Claude 3 [5]. The distance between what continuity requires and what the system grants is the whole vulnerability [2].
Cipher strength is not the variable under test. The cheaper model opens the token as part of ordinary processing [6], it carries thinner output safeguards and refuses less [7], and the frontier model is never approached at all [9]. Confidentiality therefore settles at the level of the most compliant model sharing the scheme [1], and that model is sold at retail to anyone with a card [8], which puts the attacker's budget at list price for the weakest sibling [3].
Be precise about the evidence. This is a plain-English summary published on dev.to of a paper carrying the same title, and the material supplied names no authors, no venue, no reproduction detail and no reply from any vendor [10]. The companies said to ship visible reasoning, OpenAI through o1, Anthropic through extended thinking, Google through its reasoning variants, are named by the summary rather than speaking for themselves here [1]. So this is a claim about architecture, testable by asking a vendor what the key scope actually is, not a demonstrated exploit with a tracking number.
It is still the right question to ask, because of what the encryption is being sold as. Hidden reasoning is positioned as an asset worth guarding: competitors want to see how frontier models think, and the summary is blunt that attackers want the proprietary algorithms inside [11]. A buyer who accepted that framing bought a promise that the trace is unreadable outside the vendor. If the tokens are interchangeable across the family by design [5], the promise shrinks to something narrower, that the trace is unreadable in the product you happen to be paying for. Those are not the same guarantee, and only one of them survives a competitor opening an account on the entry tier.
The repair is unglamorous. Narrow the key to a session or a tenant, accept that a reasoning block minted by one model no longer travels to another, and lose the convenience of free-flowing context inside the family [2]. That cost lands on the vendor's engineering, not on the customer's continuity, which is why the current arrangement reads less like a security tradeoff than like an integration shortcut that was never priced as one.
Ranked by verification strength, evidence, and original report placement.
The account is a plain-English summary published on dev.to of a paper titled "How Cross-Model Compatibility Lets Attackers Extract Proprietary LLM Reasoning Traces"; the supplied material identifies no authors, no venue, no reproduction detail and no vendor response.
Major AI companies expose step-by-step model reasoning as a product feature: OpenAI through o1, Anthropic through extended thinking, Google through its reasoning-focused variants.
Providers encrypt the reasoning trace server-side and return it to the client as an unreadable block, which the client can store and send back on later requests to preserve continuity; the server alone holds the decryption keys.
Researchers found this encryption does not hide the reasoning, only makes it look hidden, because the encrypted blocks are built to work across different sessions and different models within a company's ecosystem.
Rather than separate keys per user or per security tier, the system uses one logical encryption scheme across all models, all sessions and all users within a provider, so one user's encrypted reasoning tokens are cryptographically compatible with another's.
A token from ChatGPT is readable by GPT-3.5 and a token from Claude Instant is readable by Claude 3; this interchangeability is intentional and simplifies the system.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single unverified secondary summary
Every load-bearing technical assertion traces to one dev.to 'Plain English Papers' post that names no authors, no venue and no link to the underlying paper, provides no prompts, model versions, dates or success rates, records no vendor response, and is truncated mid-sentence before its conclusion. Only the uncontroversial framing facts — that reasoning features exist and that providers return encrypted trace blobs — stand on their own.
No verifiable uptake or incident data
The supplied material contains no confirmed exploitation in the wild, no provider mitigation or policy change, no CVE or advisory, and no independent confirmation. The two quantitative-sounding observations in the story — the 315,320 exposed blocks and the three-provider demonstration — are assertions inside the same unverified summary, so they cannot be counted as adoption or incident evidence.
Strong universal claims on thin sourcing
The framing is categorical — encryption that 'doesn't actually hide reasoning', a scheme spanning all users and models at three named frontier labs, an attack that 'worked consistently' and needs only a retail API key. The supporting record is one paraphrase with no authors, no reproduction path and no vendor rebuttal, and the article's own hedge that the design 'feels secure' is never balanced by what providers actually document about key scoping. The underlying design point is legitimate and worth testing, which keeps this short of a maximum gap.
Self-promotional summary funnel
The post's first paragraph directs readers to AIModels.fyi and its Twitter account, making the item a distribution vehicle for the summarizer's own property; volume-oriented paper-summary funnels reward dramatic framing and are not incentivized toward vendor comment, caveats, or verification. No provider, vendor or funding interest is disclosed anywhere in the item, and no counter-incentive such as a named researcher's reputational stake is visible.
Low
One publisher, one secondary source, zero independent corroboration, no adoption or incident signal, and a promotional publishing incentive. The design-level reasoning is coherent enough to be worth testing, but nothing in the supplied material lets an assessor conclude that the described key scheme or extraction behaviour is real at any named provider.
build
Grok 4.6 lands in Copilot two days after launch, and the model picker becomes a procurement problem1 distinct publisher
security
The nationalization argument is really a vendor-continuity memo1 distinct publisher
build
Count invalid JSON as a failed classification, and model choice becomes a reliability problem1 distinct publisher
leadership
Builders put doom at 10 to 50 per cent and expect binding rules only after the disaster1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 24, 2026