Atlas arrived Sept. 1 in early access with no price, no named partner and no paper, and the reconstruction win it publishes runs against a baseline whose own authors asked evaluators to wait. That makes it a procurement question.
Reality
- Evidence56
- Adoption8
- Hype gap+58
- Incentives82
- Confidence57
The Rust rewrite keeps pnpm 11's commands and lockfile, so adopting it costs little more than pulling from the next-12 tag, but the headline percentage describes a 1.5-second install and the cold case improved by about two thirds.
Reality
- Evidence58
- Adoption30
- Hype gap+28
- Incentives62
- Confidence55
The three track leaders have no manifest anyone can download. The curl Bench'd publishes for independent verification points at a host that does not resolve. The badge behind the board bills from $299 a month.
Reality
- Evidence42
- Adoption22
- Hype gap+18
- Incentives82
- Confidence48
A preprint proves a blind replay agent's expected score equals the source model's pass@k, which means static computer-use leaderboards have been grading environment determinism.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+16
- Incentives72
- Confidence44
Luu reports that minutes of prompting bought a 7% speedup on his own query workload. The cost curve behind it re-prices optimization backlogs and makes holdout discipline the scarce skill.
Reality
- Evidence48
- Adoption14
- Hype gap+14
- Incentives62
- Confidence55
A Metal GROUP BY kernel ran 1.7x slower than single-threaded std::unordered_map. The diagnosis came from sweeping one parameter, not from a profiler.
Reality
- Evidence62
- Adoption10
- Hype gap+12
- Incentives55
- Confidence58
SWE-Bench ProMax puts frontier coding agents on 170 curated refactoring commits. The number that should move procurement is a different one: nearly 60% of unsolved SWE-bench Verified tasks have flawed tests.
Reality
- Evidence38
- Adoption14
- Hype gap+12
- Incentives62
- Confidence42
Z.ai claims frontier agentic-coding scores at about 750B parameters, a third of Kimi K3, from extended post-training on the GLM-5.2 base. Open weights are promised in two weeks.
Publishers:interconnects.ai
Reality
- Evidence32
- Adoption24
- Hype gap+28
- Incentives68
- Confidence38