UW and Meta researchers report 59.4% on BrowseComp-Plus for a model that edits its own context, against 53.4% for Codex-style summarisation. The edited file is thrown away when the task ends, so memory that lasts across tasks is still the builder's job.
Reality
- Evidence42
- Adoption8
- Hype gap+12
- Incentives
- Insufficient
- Confidence45
One 8,192-token session on a 27B model holds 512 MB of key-value cache, or 64 KB for every token generated. How many of those sessions fit in free VRAM sets serving concurrency, and paging decides the waste.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+24
- Incentives45
- Confidence48
Anthropic screened Claude for one property, verbalizability, and found two more in the same representations. Neel Nanda reproduced the structure in open weights, and the outside commentators Anthropic invited disagree about what it is.
Reality
- Evidence40
- Adoption22
- Hype gap+12
- Incentives62
- Confidence42
NVIDIA's submission runs the same Qwen3.6-27B as the llama.cpp reference on the same Jetson board and finishes 6.4x sooner. Most of the gap comes from prompt tokens the runtime never has to prefill.
Reality
- Evidence58
- Adoption18
- Hype gap+20
- Incentives85
- Confidence62
JetBrains has taken the assembly work out of running a coding agent offline. What it could not take out is the hardware, and that is now the thing deciding who adopts.
Reality
- Evidence56
- Adoption18
- Hype gap+16
- Incentives72
- Confidence63
Muse Glimmer ships as Apache 2.0 weights sized for a 24GB card. Muse Spark 1.2 stays on Muse Code and the Meta Model API. Plan capacity for two tiers, not one.
Perspective Coverage
6 publishers
- Builder
- Builder 38%
- Operator
- Operator 31%
- Investor
- Investor 31%
Reality
- Evidence64
- Adoption48
- Hype gap+18
- Incentives78
- Confidence70