Skip to content

BuildNot yet confirmed elsewhere1 publisher2 min readPublished

Self-edited context lifts Qwen3.5-9B from 28.8% to 42.5% on BrowseComp-Plus

Meta, MIT and University of Washington researchers let models rewrite their own context, lifting a trained 9B model from 28.8% to 42.5% on BrowseComp-Plus. The gains come from specified benchmark tasks, and rewriting earlier context invalidates the model cache, so we would keep hand-built compaction pipelines as the production default.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Self-edited context lifts Qwen3.5-9B from 28.8% to 42.5% on BrowseComp-Plus
Generated illustration

What happened

  • The models treat their context as a file they can rewrite without restrictions, in place of the compaction, offloading and retrieval rules a harness normally applies.
  • With no training, the models scored 11.4% higher on BrowseComp-Plus using 21.5% fewer FLOPs, and 5% higher on 12-hour EdgeBench using 59% fewer FLOPs.
  • An in-context learning variant raised ContextBench accuracy by up to 35.9 percentage points while using less compute.
  • The reinforcement-learning run behind the 9B result used 12% fewer FLOPs, with compute efficiency as a secondary training objective after task success.

Why it matters

  • cost Savings reported in FLOPs can be offset by recompute on serving stacks that rely on cached prefixes, and the team running inference pays for each invalidated cache.
  • exposure Anything the model chooses to keep in its own context outlives the turn it arrived in, so security review of an agent has to cover the context file as well as the prompt.
  • constraint A fact the model deletes cannot be retrieved later, so the approach carries more risk on tasks where one dropped detail can sink the result.

In this design the tuning work moves into the prompt. The in-context learning variant gets its context strategy as a natural-language instruction. It then refines that strategy through an iterative skill-optimization loop [6]. Users can steer context management by telling the agent which strategy they want, and the model can evolve a skill document of context-management procedures to reuse later [5]. For a team that maintains a compaction pipeline, the thing under review becomes a prompt plus a document the model edits itself. The researchers say learned strategies may surpass "existing human priors" [19], a polite name for the predefined rules compaction usually runs on [3].

The reinforcement-learning figure is internally consistent. 42.5 minus 28.8 is 13.7 points, and 13.7 over 28.8 gives the 47.6% relative gain InfoQ reports [18]. The InfoQ summary does not say which baseline the zero-shot gains are measured against, or how the FLOPs counts treat cache misses. For the numbers to transfer, a team's tasks have to resemble the benchmarks, two of which run for 12 and 24 hours [7][8]. On the 24-hour multi-repository agent-swarm task, zero-shot CLMs reported 65% greater improvement at the same compute [8]. @Rennix7t wrote on X that "the accuracy and computational power results in the paper come from specified tasks" [14].

Serving cost is where we would push hardest. Combinatorilliance argued on Reddit that the approach is not production-ready because rewriting the context invalidates the model cache [16]. The researchers' caching trick "isn't ideal either and has some drawbacks that need to be taken into consideration," the same commenter wrote [17]. We'd expect the penalty to grow with how early an edit lands, since everything after the changed message has to be processed again. We think that favours fixed rules on cost. A compaction rule changes the prefix only when it fires. A model allowed to rewrite old messages [4] can change it on any turn.

The researchers flag the security cost themselves. The editable context "can become another channel through which prompt injections or self-generated instructions persist across turns," they wrote [12]. They also concede that more control over context does not necessarily produce better behavior, because a model can still choose badly what to keep, change or discard [13]. We think hand-built context pipelines stay the right production default on this evidence. They are also the control group a team needs to tell whether a model-managed context saves more compute than its cache misses cost on real work. @omarsar0 called the approach an interesting research direction but voiced reservations about trusting a model to manage its context end-to-end [15].

What to watch

  • Serving-cost measurements that count prompt-cache misses, showing whether the FLOPs savings hold on production inference stacks.
  • Independent runs of model-managed context on workloads outside BrowseComp-Plus, EdgeBench and ContextBench.
  • Red-team results on injected instructions that a model retains in its own context across turns.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence45
Adoption
Insufficient
Hype gap+20
Incentives
Insufficient
Confidence40
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Researchers from Meta, MIT, and the University of Washington introduced Context Language Models (CLMs), which let language models manage and edit their own context instead of relying on predefined mechanisms for summarization, compression and information retrieval.

    ReportedSupportedSource: InfoQView cited source
  2. [2]

    The researchers implemented self-managed context by treating the context as a file the model can update with no restrictions, as an alternative to conventional strategies such as compaction, offloading and retrieval.

    ReportedSupportedSource: InfoQView cited source
  3. [3]

    Summarization can discard critical details or introduce inaccuracies, compaction often uses a predefined set of rules, and external memory requires agents to decide what to bring back into the context.

    ReportedSupportedSource: InfoQView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. infoq.com

    1 article · October 11, 2026

    Context Language Models: Self-Managing Context to Improve Performance and Reduce Compute Costs

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Loading related stories