The institute says an average pass rate scores a refusal and an inability the same way, so it read almost 6,400 evaluation transcripts to see which was happening. Pass rates stay in its pre-deployment reports.
Publishers:aisi.gov.uk
Reality
- Evidence62
- Adoption32
- Hype gap0
- Incentives45
- Confidence58
buildOne report1 publisher tau-bench scores agents on the database state they leave behind and then reruns each task eight times, and the consistency figure that falls out is a better launch gate than the single-trial accuracy most teams quote.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+8
- Incentives55
- Confidence58
buildOne report1 publisher AgentArch sweeps orchestration, ReAct versus function calling, memory scope and a thinking tool across 18 setups on frontier models. Because the best cell moves with the model, the grid is what you reuse and the ceiling is what you budget for.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+8
- Incentives55
- Confidence50
buildOne report1 publisher A dev.to post puts the standard-versus-agentic line at when the retrieval decision gets made, and the interesting part is what a fixed top-k pipeline gives up the moment a planner starts writing plans per question.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+12
- Incentives20
- Confidence42
buildOne report1 publisher latent.space argues models and harnesses improved together, and that models keep swallowing the harness. If so, most scaffolding you write is scheduled for deletion. Permissions are not.
Reality
- Evidence42
- Adoption48
- Hype gap+18
- Incentives
- Insufficient
- Confidence38