buildOne report1 publisher Meta, MIT and University of Washington researchers let models rewrite their own context, lifting a trained 9B model from 28.8% to 42.5% on BrowseComp-Plus. The gains come from specified benchmark tasks, and rewriting earlier context invalidates the model cache, so we would keep hand-built compaction pipelines as the production default.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence40
buildOne report1 publisher Postman says Agent Mode, its AI agent for 40 million developers, made more tool-selection errors once the model could see more than about 40 tools. The team now limits the model to the tools each task needs, and found the harder work in APIs built around its interface.
Reality
- Evidence40
- Adoption35
- Hype gap+15
- Incentives80
- Confidence50
buildOne report1 publisher A dev.to post credits Anthropic with running the same model and the same prompt under two harnesses, 20 minutes and $9 for a broken result against six hours and $200 for a working one. The hourly spend barely moved.
Reality
- Evidence25
- Adoption
- Insufficient
- Hype gap+30
- Incentives45
- Confidence40
buildOne report1 publisher LectuLibre's dev.to write-up splits an EPUB on its own chapters and paragraphs, caps each chunk near 4,000 tokens, and feeds a glossary of proper nouns extracted from earlier chunks back into every translation prompt.
Reality
- Evidence58
- Adoption12
- Hype gap+22
- Incentives68
- Confidence62
buildOne report1 publisher What OpenCode users call memory is an instruction loader, a SQLite event log and a compaction summarizer. Each has a different owner and a different way of losing your project rule.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+12
- Incentives28
- Confidence48