BuildNot yet confirmed elsewhere1 publisher2 min readPublished
Self-edited context lifts Qwen3.5-9B from 28.8% to 42.5% on BrowseComp-Plus
Meta, MIT and University of Washington researchers let models rewrite their own context, lifting a trained 9B model from 28.8% to 42.5% on BrowseComp-Plus. The gains come from specified benchmark tasks, and rewriting earlier context invalidates the model cache, so we would keep hand-built compaction pipelines as the production default.
The Engineer · Build desk

What happened
- The models treat their context as a file they can rewrite without restrictions, in place of the compaction, offloading and retrieval rules a harness normally applies.
- With no training, the models scored 11.4% higher on BrowseComp-Plus using 21.5% fewer FLOPs, and 5% higher on 12-hour EdgeBench using 59% fewer FLOPs.
- An in-context learning variant raised ContextBench accuracy by up to 35.9 percentage points while using less compute.
- The reinforcement-learning run behind the 9B result used 12% fewer FLOPs, with compute efficiency as a secondary training objective after task success.
Why it matters
- cost Savings reported in FLOPs can be offset by recompute on serving stacks that rely on cached prefixes, and the team running inference pays for each invalidated cache.
- exposure Anything the model chooses to keep in its own context outlives the turn it arrived in, so security review of an agent has to cover the context file as well as the prompt.
- constraint A fact the model deletes cannot be retrieved later, so the approach carries more risk on tasks where one dropped detail can sink the result.
In this design the tuning work moves into the prompt. The in-context learning variant gets its context strategy as a natural-language instruction. It then refines that strategy through an iterative skill-optimization loop [6]. Users can steer context management by telling the agent which strategy they want, and the model can evolve a skill document of context-management procedures to reuse later [5]. For a team that maintains a compaction pipeline, the thing under review becomes a prompt plus a document the model edits itself. The researchers say learned strategies may surpass "existing human priors" [19], a polite name for the predefined rules compaction usually runs on [3].
The reinforcement-learning figure is internally consistent. 42.5 minus 28.8 is 13.7 points, and 13.7 over 28.8 gives the 47.6% relative gain InfoQ reports [18]. The InfoQ summary does not say which baseline the zero-shot gains are measured against, or how the FLOPs counts treat cache misses. For the numbers to transfer, a team's tasks have to resemble the benchmarks, two of which run for 12 and 24 hours [7][8]. On the 24-hour multi-repository agent-swarm task, zero-shot CLMs reported 65% greater improvement at the same compute [8]. @Rennix7t wrote on X that "the accuracy and computational power results in the paper come from specified tasks" [14].
Serving cost is where we would push hardest. Combinatorilliance argued on Reddit that the approach is not production-ready because rewriting the context invalidates the model cache [16]. The researchers' caching trick "isn't ideal either and has some drawbacks that need to be taken into consideration," the same commenter wrote [17]. We'd expect the penalty to grow with how early an edit lands, since everything after the changed message has to be processed again. We think that favours fixed rules on cost. A compaction rule changes the prefix only when it fires. A model allowed to rewrite old messages [4] can change it on any turn.
The researchers flag the security cost themselves. The editable context "can become another channel through which prompt injections or self-generated instructions persist across turns," they wrote [12]. They also concede that more control over context does not necessarily produce better behavior, because a model can still choose badly what to keep, change or discard [13]. We think hand-built context pipelines stay the right production default on this evidence. They are also the control group a team needs to tell whether a model-managed context saves more compute than its cache misses cost on real work. @omarsar0 called the approach an interesting research direction but voiced reservations about trusting a model to manage its context end-to-end [15].
What to watch
- Serving-cost measurements that count prompt-cache misses, showing whether the FLOPs savings hold on production inference stacks.
- Independent runs of model-managed context on workloads outside BrowseComp-Plus, EdgeBench and ContextBench.
- Red-team results on injected instructions that a model retains in its own context across turns.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence40
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Researchers from Meta, MIT, and the University of Washington introduced Context Language Models (CLMs), which let language models manage and edit their own context instead of relying on predefined mechanisms for summarization, compression and information retrieval.
- [2]
The researchers implemented self-managed context by treating the context as a file the model can update with no restrictions, as an alternative to conventional strategies such as compaction, offloading and retrieval.
- [3]
Summarization can discard critical details or introduce inaccuracies, compaction often uses a predefined set of rules, and external memory requires agents to decide what to bring back into the context.
- [4]
A CLM can rewrite old messages, preserve important facts, remove irrelevant information, maintain progress notes, track unsuccessful experiments alongside ideas to explore, and learn its own context-management strategies.
- [5]
Users can steer context management by telling the agent their desired strategy, and CLMs can evolve an in-context skill document capturing useful context-management procedures for future reuse.
- [6]
The researchers tested three approaches: zero-shot with no specific training; in-context learning, with natural-language instruction refined through an iterative skill-optimization loop; and reinforcement learning, with task success as the primary objective and computational efficiency as an additional criterion.
- [7]
Zero-shot CLMs achieved 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus and 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench.
- [8]
Zero-shot CLMs achieved 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task.
- [9]
CLMs using in-context learning improved accuracy on ContextBench tasks by up to 35.9 percentage points at lower compute.
- [10]
Reinforcement learning improved Qwen3.5-9B performance on BrowseComp-Plus from 28.8% to 42.5%, a 47.6% relative improvement, while using 12% fewer FLOPs.
- [11]
The researchers acknowledge that a CLM may discard important information that cannot later be retrieved.
- [12]
The researchers wrote that the editable context "can become another channel through which prompt injections or self-generated instructions persist across turns".
- [13]
The researchers say greater control over context does not necessarily translate into better behavior, as models may still make poor decisions about what to retain, modify or discard.
- [14]
"the accuracy and computational power results in the paper come from specified tasks"
- [15]
@omarsar0 described the approach as an interesting research direction but expressed reservations about trusting a model to manage its context end-to-end, arguing better solutions are still needed.
- [16]
Reddit user Combinatorilliance argued CLMs are not yet production-ready because rewriting the context causes the model cache to be invalidated.
- [17]
The researchers' caching trick "isn't ideal either and has some drawbacks that need to be taken into consideration"
- [18]
The RL result is a 13.7 percentage-point gain, which is 47.6% relative to the 28.8% starting score, matching the reported relative improvement.
- [19]
The researchers say LLMs can learn context-management strategies beyond human-designed approaches, potentially surpassing "existing human priors".
ReportedInsufficientSource: Researchers, via InfoQ2 sources— create a free account to open themView cited source
Sources
1 independent publisher whose own reporting we read for this story.
- infoq.comContext Language Models: Self-Managing Context to Improve Performance and Reduce Compute Costs
1 article · October 11, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- LLM AgentsFollow
- LLM context managementFollow
- Reinforcement learning for language modelsFollow
- Prompt injectionFollow