One developer's self-hosted Langfuse traced 91 of 1,387 agent model calls in 14 days because only one of ten profiles had its keys. The plugin fails open by design, so nine untraced profiles, one a coding agent with more than five times the traced calls, raised no error.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap0
- Incentives
- Insufficient
- Confidence50
Researchers writing on LessWrong found only Nemotron 3 Super and Qwen3 32B changed refusals when test suspicion was removed from their reasoning traces. In most other models, test talk in a trace looks like general caution, so it is weak evidence of eval gaming.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap0
- Incentives
- Insufficient
- Confidence45
One developer's audit of Claude Code found 97 of 200 agent steps could run on a local model, against 4 of 100 whole requests. That makes the agent step the unit to route on, on evidence from one person's sessions and one RTX 4070.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives25
- Confidence35
Jared Palmer's Kev now ships as 0.8B, 4B and 9B variants that score typed answer options over frozen Qwen3.5 weights, and on his own development comparison hosted Jev still scores higher than the largest of them.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+10
- Incentives35
- Confidence60
NVIDIA's SWE-Serve scores the same 627 patches twice on 19 SGLang tasks, once with the live-serving tests and once without. The pass rate falls from 69.4% to 45.9%, and 242 of the 276 live tests came from SGLang itself.
Reality
- Evidence58
- Adoption30
- Hype gap−10
- Incentives70
- Confidence60
Bespoke Nimble's LoRA fine-tune of Qwen3.5-9B scored 90% against Jev's 93% on an eval it curated itself, and five more replications landed at scales between 421M and 35B parameters. No standard benchmark exists for the category yet.
Reality
- Evidence32
- Adoption46
- Hype gap+42
- Incentives72
- Confidence38
Per-tensor layout maps now drive his GGUF releases. The sensitivity data behind them came out of more than 1,000 quantizations of two Qwen3.5 models, scored by KL divergence against bf16 on wikitext-2-raw.
Reality
- Evidence57
- Adoption28
- Hype gap+8
- Incentives42
- Confidence55
Publishing weights and publishing something a team can actually deploy are different acts, and N2.5 lands on both sides of that line depending on which tier you pick and whether its weights exist yet.
Reality
- Evidence45
- Adoption15
- Hype gap+30
- Incentives68
- Confidence50
p-for-llm puts a 29-expert, top-1-routed mixture of experts on an ESP32-P4 and got two points on Hacker News. The routing, not the quantization, is the load-bearing part.
Reality
- Evidence28
- Adoption8
- Hype gap+18
- Incentives45
- Confidence34