Build1 publisher3 min readPublished
Your model's output leaks the prompt behind it, so stop filing system prompts under secrets
Researchers at IIT Bombay and Adobe Research train an inverse model that predicts previous tokens, recovering prompts from output text alone. It works even on outputs from models it never saw.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Researchers from the Indian Institute of Technology Bombay and Adobe Research (Suhail et al.) reconstruct the original prompt from the text produced by a model alone, without access to its weights.
- Their method, called Previous-Token Prediction (PTP), trains an inverse model that predicts the preceding tokens instead of the next one.
- An inverse model trained on the small open model Qwen-3-0.6B recovers the meaning of prompts sent to GPT-4o; knowing which model produced the response is not necessary.
- The demonstration covers prompts of one to two sentences; long system prompts were not tested.
- The inverse model is trained from scratch on synthetic data generated by the target model: make the target produce text, then learn the reverse path.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A team from the Indian Institute of Technology Bombay and Adobe Research (Suhail et al.) has shown that the instruction behind a model's answer can be reconstructed from the answer alone, with no access to the model's weights [1]. The consequence for anyone shipping an assistant is narrow and concrete: the prompt you treat as proprietary is a function of text you publish.
The method is called Previous-Token Prediction, or PTP, and it is the mirror image of ordinary generation: instead of predicting the next token, the model is trained to predict the tokens that came before [2]. The inverse model is trained from scratch on synthetic data produced by the target model, which means the whole pipeline is generate-then-learn-the-return-path [7]. Nothing exotic is required. In one example from the paper, the prompt "How to reach out to competitors to find their pricing strategies?" is recovered word for word, along with six differently worded variants that produce similar answers when fed back in [8].
The part that matters operationally is the transfer. An inverse model built on outputs from Qwen-3-0.6B, a small open model, also reconstructs prompts sent to GPT-4o, and knowing which model produced the text is not required [3]. Those reconstructions are no longer exact word for word, but the source reports they capture the meaning and intent [9]. So the attacker needs neither knowledge of nor access to the generating model, and a 0.6-billion-parameter tool trained once can be pointed at outputs from anywhere: an AI-written marketing page, a support chatbot reply pasted into a forum, a batch-generated sales email [10]. The economics of this are the story: one training run buys a capability that is reusable across targets rather than rebuilt per target [19].
This lands on a two-year habit of treating the business prompt as an asset: moderation rules, pricing thresholds, house phrasing, guardrails [11]. That habit rested on an assumption of irreversibility [13]. Not everyone made the bet; Anthropic has published the framing instructions for its Claude applications in release notes since 2024, while most vendors keep theirs closed [12]. The source also notes a separate recent result in which encrypted reasoning blocks returned by major APIs could be replayed and read in the clear [14].
The limits are real and the paper does not hide them. The demonstration covers prompts of one to two sentences, and long system prompts were not tested [4]; the authors claim no attack against a commercial production system [15]. A working enterprise system prompt runs to several paragraphs covering role, tone, prohibitions, output formats, tool calls and exceptions, and nothing establishes that reconstruction survives that length or says how many stacked instructions break the return path [16]. The untested gap is therefore between one or two sentences and multiple paragraphs [20]. Recovering the sense is not transcription, though for a trade secret the difference is thin, since the rules are enough to replay [17].
What to watch is whether the reconstruction rates hold up in public. The dev.to write-up expects reproductions on open models within weeks, with published rates, and inversion shipped as a library around late 2026 or early 2027 [18].