Microsoft Research Asia open-sourced Agent Lightning v1.0, a 3,500-line framework whose recipe lifted a 9B model's SWE-bench Verified score by 14.6 points. It runs reinforcement learning on the agent a team already deploys, sitting as a proxy between that agent and its model.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+15
- Incentives60
- Confidence40
build1 publisherOne report Aura Memory's developer kept almost all memory value on 120 LongMemEval questions by storing only the user's words, at an eighth of the size. Summaries in the same tests dropped facts and invented advice, so raw messages are the safer default.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+10
- Incentives40
- Confidence45
build1 publisherOne report A study of eight frontier models on SWE-bench Verified puts agentic coding at roughly 1,000 times the token cost of code chat, dominated by input, with the models' own pre-run estimates correlating no better than 0.39.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+15
- Incentives30
- Confidence55
build1 publisherOne report NVIDIA's developer blog sets out the five-level rollup from step to benchmark and the two scores that read the same trace. Step-level says where the chain broke. End-to-end reads the environment and says whether the refund posted.
Reality
- Evidence55
- Adoption20
- Hype gap+12
- Incentives50
- Confidence60
build1 publisherOne report The study counts only complexity and dead code, because those are the two pyscn metrics that stay exact when you analyze just the files a patch touched. The human's own commit trips the same rule 24% of the time.
Reality
- Evidence55
- Adoption20
- Hype gap+12
- Incentives62
- Confidence60
build1 publisherOne report OpenHands logged a conversation stuck for 8 hours 21 minutes while /health kept returning 200. CodeFlowMu logged a QA release stamped 17 seconds before the upstream report it depended on.
Reality
- Evidence42
- Adoption18
- Hype gap+8
- Incentives62
- Confidence47
Token traffic, survey reach and production-model ledgers rank different vendors because they count different things. The autonomy figures say the hard part is still unbought.
Reality
- Evidence58
- Adoption64
- Hype gap+32
- Incentives68
- Confidence52
build1 publisherOne report Hermes Agent treats the runtime as the durable asset and the model as a swappable input. The lock-in moves into your own repo, and the administrator's job moves with it.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+32
- Incentives74
- Confidence42