Build1 distinct publisher3 min readUpdated
Researchers at IIT Bombay and Adobe Research train an inverse model that predicts previous tokens, recovering prompts from output text alone. It works even on outputs from models it never saw.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
A team from the Indian Institute of Technology Bombay and Adobe Research (Suhail et al.) has shown that the instruction behind a model's answer can be reconstructed from the answer alone, with no access to the model's weights [1]. The consequence for anyone shipping an assistant is narrow and concrete: the prompt you treat as proprietary is a function of text you publish.
The method is called Previous-Token Prediction, or PTP, and it is the mirror image of ordinary generation: instead of predicting the next token, the model is trained to predict the tokens that came before [2]. The inverse model is trained from scratch on synthetic data produced by the target model, which means the whole pipeline is generate-then-learn-the-return-path [7]. Nothing exotic is required. In one example from the paper, the prompt "How to reach out to competitors to find their pricing strategies?" is recovered word for word, along with six differently worded variants that produce similar answers when fed back in [8].
The part that matters operationally is the transfer. An inverse model built on outputs from Qwen-3-0.6B, a small open model, also reconstructs prompts sent to GPT-4o, and knowing which model produced the text is not required [3]. Those reconstructions are no longer exact word for word, but the source reports they capture the meaning and intent [9]. So the attacker needs neither knowledge of nor access to the generating model, and a 0.6-billion-parameter tool trained once can be pointed at outputs from anywhere: an AI-written marketing page, a support chatbot reply pasted into a forum, a batch-generated sales email [10]. The economics of this are the story: one training run buys a capability that is reusable across targets rather than rebuilt per target [19].
This lands on a two-year habit of treating the business prompt as an asset: moderation rules, pricing thresholds, house phrasing, guardrails [11]. That habit rested on an assumption of irreversibility [13]. Not everyone made the bet; Anthropic has published the framing instructions for its Claude applications in release notes since 2024, while most vendors keep theirs closed [12]. The source also notes a separate recent result in which encrypted reasoning blocks returned by major APIs could be replayed and read in the clear [14].
The limits are real and the paper does not hide them. The demonstration covers prompts of one to two sentences, and long system prompts were not tested [4]; the authors claim no attack against a commercial production system [15]. A working enterprise system prompt runs to several paragraphs covering role, tone, prohibitions, output formats, tool calls and exceptions, and nothing establishes that reconstruction survives that length or says how many stacked instructions break the return path [16]. The untested gap is therefore between one or two sentences and multiple paragraphs [20]. Recovering the sense is not transcription, though for a trade secret the difference is thin, since the rules are enough to replay [17].
What to watch is whether the reconstruction rates hold up in public. The dev.to write-up expects reproductions on open models within weeks, with published rates, and inversion shipped as a library around late 2026 or early 2027 [18].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Researchers from the Indian Institute of Technology Bombay and Adobe Research (Suhail et al.) reconstruct the original prompt from the text produced by a model alone, without access to its weights.
Their method, called Previous-Token Prediction (PTP), trains an inverse model that predicts the preceding tokens instead of the next one.
An inverse model trained on the small open model Qwen-3-0.6B recovers the meaning of prompts sent to GPT-4o; knowing which model produced the response is not necessary.
The demonstration covers prompts of one to two sentences; long system prompts were not tested.
The inverse model is trained from scratch on synthetic data generated by the target model: make the target produce text, then learn the reverse path.
For GPT-4o outputs the reconstructions are no longer word-exact, but they capture the meaning and the intent.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Thin: one secondary write-up, no primary citation
Everything rests on a single French-language dev.to post summarizing an unlinked paper by Suhail et al. (IIT Bombay / Adobe Research). The mechanism is described coherently and one qualitative example is given, but no reconstruction rates, datasets, baselines, venue or paper link appear, and no independent reproduction is cited. The source's own scope disclosures (short prompts only, no production attack) are credible but limit what the evidence can carry.
No adoption signal in supplied sources
The cluster contains no release, deployment, benchmark result, security incident, pricing or licensing event, and no usage disclosure. The write-up states the authors claim no attack on a production system, and the packaging of inversion as a library is presented as a forecast, not an observed shipment. Nothing in the supplied material measures uptake.
Overstated relative to demonstrated scope
The cluster headline tells readers to stop filing system prompts under secrets, and the write-up generalizes to enterprise moderation rules, pricing thresholds and guardrails. The demonstrated result is narrower: one-to-two-sentence prompts, word-exact recovery only on a small open model, meaning-level recovery on GPT-4o, no production system attacked, and no published reconstruction rates. The source itself concedes the whole scenario hinges on an untested variable — the prompt length the method tolerates — which is the gap between claim and demonstration.
Modest, mostly disclosed
The disclosed incentive structure is limited: an academic group plus a corporate research lab (Adobe Research) publishing an offensive-capability result, and a personal developer-blog author who supplies a dated forecast and a prescriptive hygiene checklist that reward attention. No funding relationship, vendor sponsorship, product, or commercial disclosure appears in the supplied material, and the source does not sell a mitigation, so scoring stays low.
Low
Direction is plausible and internally consistent, and the source is unusually candid about its limits, which raises trust in the scope statements. But one publisher, no primary paper reference, no quantitative results, an uncited companion result on encrypted reasoning blocks, and zero adoption evidence keep overall confidence low. The claim class (prompt inversion from output text) is actionable as hygiene regardless; the claim magnitude is not yet verifiable.
build
The Tokenizer Is Your Real Price List, Not the Per-Million Rate Card1 distinct publisher
build
An OAuth login now lets Claude rewrite, or delete, your live ElevenLabs voice agent1 distinct publisher
science
Text watermarks land on 2 December. The detection they imply does not.1 distinct publisher
build
The $559M-versus-$12.3B quarter matters more than the $65B run rate4 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 18, 2026