Skip to content

Build1 publisher2 min readPublished

UW and Meta's Context Language Models beat Codex-style summarisation by six points on BrowseComp-Plus

UW and Meta researchers report 59.4% on BrowseComp-Plus for a model that edits its own context, against 53.4% for Codex-style summarisation. The edited file is thrown away when the task ends, so memory that lasts across tasks is still the builder's job.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying UW and Meta's Context Language Models beat Codex-style summarisation by six points on BrowseComp-Plus
Generated illustration

What happened

  • Rulin Shao and colleagues at the University of Washington and Meta Superintelligence Labs published the paper on 29 September 2026, with co-authors at MIT and Trillium Labs.
  • Trained with stepwise GRPO, Qwen3.5-9B rose from 28.8% to 42.5% on BrowseComp-Plus and matched a summarisation baseline trained the same way while using 38.8% fewer FLOPs.
  • On 12-hour EdgeBench runs the method scored 44.6 against 42.3 for the strongest baseline, using 59% fewer FLOPs.
  • The team put the harness, the in-context and RL modules and a serving optimisation called Suffix Cache Reuse on GitHub, but released no trained weights.
  • The paper's own safety section warns that the write access the method depends on also lets a model plant instructions for itself.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure A team that saves the context file as memory for later tasks would also save any instructions the model planted for itself, so that file needs checking before it is stored.
  • constraint A commercial team cannot ship the released harness under CC BY-NC 4.0 and has to rebuild the file-sync loop itself to use the method.
  • capability The prompted route needs only a harness and a skill document, so a team can test it on the model it already calls without training anything.

A Context Language Model keeps its transformer unchanged [20] and puts the change in the harness. Context stops being an append-only transcript and becomes a file the model can edit anywhere with ordinary Bash commands [21][7]. The authors wrote that "edits to the context file are automatically synchronized with the LM's context and sent to the LLM server" [8]. Generation then continues on the edited version [8]. A model can compress twenty tool calls into two lines [9]. It can also delete a stale search result, rewrite its own plan, or keep a running scoreboard at the top of the file and update it in place [9].

On the rounded scores, the prompted Qwen3.6-27B result is 6.0 points ahead of Codex-style summarisation [1]. That is 11.2% relative [2]. The write-up's table gives 11.4% [3]. The compute saving was 21.5% fewer FLOPs, measured at a 32K context limit [2]. I would expect that saving to carry over only to agents that hit a similar ceiling during long search loops.

The authors evolved the in-context skill document [10]. The agent produced rollouts, a proposer model drafted candidate skills from those traces, and a development split picked the winner [10]. That route needs no training run. Reinforcement learning is the other route, and the 9B result comes from a stepwise GRPO run [11].

The baselines include Codex-style summaries, MEM1, Self-Compact, context folding, Recursive Language Models, Mini-SWE-Agent and an agentic context management method from Li et al. [15]. In a six-agent Software World run lasting 24 hours, CLM agents produced a 65% greater downstream speedup than summary-based swarms at the same compute [13]. The largest gap, up to 35.9 points over prior strategies, came on ContextBench's KV Store task [14]. ContextBench is listed as "coming soon" [18]. Until it ships, nobody outside the team can check that figure.

I think the design is right for work inside a single task. According to the paper, it beat the hand-written context rules the authors tested it against [22]. The only thing it asks of the serving side is to take an edited context and keep generating [8]. The code ships under CC BY-NC 4.0, and the write-up says that rules out commercial use [19].

What to watch

  • The public release of ContextBench, which would let outside teams test the claimed 35.9-point KV Store gain.
  • Whether the authors release trained weights or relicense the code under terms that allow commercial use.
  • Independent runs of the prompted method on models other than Qwen3.6-27B, or at context limits above 32K.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories