Cua's pitch is an OS-level driver that delivers clicks and keystrokes in the background on macOS, Windows and Linux. The post hedges it with "where supported by the platform". A team has to test that clause before designing around it.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+40
- Incentives
- Insufficient
- Confidence32
Researchers held a coding harness's execution loop fixed and varied planning, tools and context management across four models, and found that what each component is worth depends on how strong the model already is.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+12
- Incentives55
- Confidence58
The DeepDeck benchmark runs one task twice, once with the WebMCP bundle installed and once with the tool list forced empty, freezing the model configuration in between and leaving correctness unscored unless you supply an answer you checked yourself.
Reality
- Evidence45
- Adoption15
- Hype gap−10
- Incentives35
- Confidence55
Arize and Fireworks priced ten models on what a completed command-line task costs, across 2,400 runs. The winner on that metric is an open model with the worst pass rate in the study and the thinnest coverage.
Publishers:arize.com
Reality
- Evidence62
- Adoption18
- Hype gap+14
- Incentives75
- Confidence55
CRMArena-Pro reports about 58% single-turn success and 35% multi-turn over nineteen expert-validated tasks, and the per-skill breakdown inside those averages is the part that should decide where an agent gets pointed.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap−12
- Incentives52
- Confidence58
JetBrains traced a model through fifteen C# refactoring tasks and found it simulating structure with sed, git and the compiler. Wiring in Rider's real engine cut median time by 83%.
Reality
- Evidence58
- Adoption30
- Hype gap+20
- Incentives84
- Confidence55
Arize and Fireworks ran ten models against 40 agent tasks and found the cheapest model per finished job also had the worst pass rate. Coverage, not price, is the binding constraint.
Publishers:arize.com
Reality
- Evidence52
- Adoption20
- Hype gap+22
- Incentives78
- Confidence45
A controlled evaluation across four benchmarks found centralized coordination lifted financial reasoning 80.9%, while every multi-agent variant tested made strict sequential planning worse.
Reality
- Evidence46
- Adoption18
- Hype gap+12
- Incentives62
- Confidence52
An arXiv evaluation across OpenAI, Anthropic and Google reports that naive full-context caching can raise latency, while excluding dynamic tool results gives more consistent gains.
Reality
- Evidence64
- Adoption22
- Hype gap+14
- Incentives42
- Confidence58
Sol-Luna's supervisor could have split independent modules across parallel workers. Given a free choice across six benchmark runs, it kept the work for itself every time.
Reality
- Evidence46
- Adoption
- Insufficient
- Hype gap−12
- Incentives42
- Confidence43