buildConfirmed2 publishers Microsoft has put Decision-1, a model that scores fixed choices, in Foundry at $0.042 per million input tokens. Agent teams now have to test whether its scores, so far benchmarked only by Microsoft, are calibrated well enough to decide when a case goes to a human.
Reality
- Evidence55
- Adoption15
- Hype gap+25
- Incentives70
- Confidence60
Microsoft released Decision-1, a model that scores fixed answer options and, by its own benchmarks, runs 35 times faster than GPT-6 Sol. Whether routing work moves off general-purpose models now depends on independent tests and on pricing that confirm those self-reported figures.
Reality
- Evidence30
- Adoption15
- Hype gap+35
- Incentives65
- Confidence35
Microsoft Research Asia open-sourced Agent Lightning v1.0, a 3,500-line framework whose recipe lifted a 9B model's SWE-bench Verified score by 14.6 points. It runs reinforcement learning on the agent a team already deploys, sitting as a proxy between that agent and its model.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+15
- Incentives60
- Confidence40
buildOne report1 publisher UW and Meta researchers report 59.4% on BrowseComp-Plus for a model that edits its own context, against 53.4% for Codex-style summarisation. The edited file is thrown away when the task ends, so memory that lasts across tasks is still the builder's job.
Reality
- Evidence42
- Adoption8
- Hype gap+12
- Incentives
- Insufficient
- Confidence45
buildOne report1 publisher Ollama 0.35 adds a /v1/systemone endpoint returning a choice, yes/no or score from models run on the device. Its 9B Nimble model matched human moderation labels 70.3% of the time in its maker's test, so each team still sets its own review threshold.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives60
- Confidence45
buildOne report1 publisher Probes on Qwen3.5-27B reading other models' text came within 0.004 AUROC, on average, of probing authors up to 397B directly, a LessWrong post reports. Every tested pair was open-weight, so auditors who apply the method to closed models get the reader's view of the text and cannot measure that gap.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence30
buildOne report1 publisher A dev.to guide to running local models on 8GB prices the KV cache between 15KB and 160KB per token depending on architecture. At 32K tokens held, that spread is the difference between 0.5GB and 5GB of a fixed budget.
Reality
- Evidence22
- Adoption
- Insufficient
- Hype gap+45
- Incentives62
- Confidence58
The band it joined is one where the best closed models still finish under 10% of tasks end to end at roughly $50 and 20 minutes each, so what post-training buys a law firm is price and hosting rather than capability.
Publishers:harvey.ai
Reality
- Evidence41
- Adoption
- Insufficient
- Hype gap+32
- Incentives76
- Confidence46
buildOne report1 publisher A June finding said local models cannot iterate on code. It said so about chat UIs, and it named the fix. July supplied that fix and measured it. The follow-up test result is the number worth reading.
Reality
- Evidence46
- Adoption12
- Hype gap+12
- Incentives32
- Confidence44
buildOne report1 publisher Microsoft reports Qwen3.5-9B going from 41.8% to 56.4% on SWE-bench Verified using 6K examples. The load-bearing change is who owns the agent loop, and what the service boundary costs.
Reality
- Evidence32
- Adoption18
- Hype gap+28
- Incentives66
- Confidence44
buildOne report1 publisher One model was verified at 30.16 in the official harness and reported at 100.00 in NVIDIA's. Microsoft's Agent Lightning now trains the harness into the weights.
Reality
- Evidence54
- Adoption58
- Hype gap+28
- Incentives66
- Confidence48
buildOne report1 publisher PlannerCritic's sweeps are cheap enough to gate every release. The harder problem is telling a clean run from a rig that quietly counted nothing.
Reality
- Evidence34
- Adoption11
- Hype gap+24
- Incentives68
- Confidence33