build1 publisherOne report Microsoft Research's Webwright lifted GPT-5.4 from 33.5% to 60.1% on 200 long web tasks by having it write browsing code from a terminal. Because the baseline steered by screen coordinates, the result compares writing code with clicking pixels.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives60
- Confidence45
Berkeley's RDI center attacked the step where each benchmark computes its score, and without solving a task its own scorecard reports 100% on five of the eight, about 98% on GAIA and 73% on OSWorld.
Publishers:rdi.berkeley.edu
Reality
- Evidence45
- Adoption35
- Hype gap+20
- Incentives55
- Confidence50
build1 publisherOne report The mitigation vendors cite when asked about prompt injection is instruction hierarchy. In WASP's sandbox, agents built on models that have it still began following text a human wrote into a webpage.
Reality
- Evidence64
- Adoption
- Insufficient
- Hype gap+12
- Incentives45
- Confidence58
build1 publisherOne report A dev.to write-up puts agent-browser inside a UI evaluation harness so a coding agent hunts for the Settings dialog through an accessibility snapshot, with a one-line Playwright assertion deciding pass or fail.
Reality
- Evidence35
- Adoption20
- Hype gap+10
- Incentives25
- Confidence45
build1 publisherOne report A new diagnostic benchmark treats the execution layer as something to vary rather than a fixed backdrop. Its conclusion is that agent capability belongs to a model-harness pair.
Reality
- Evidence42
- Adoption10
- Hype gap+22
- Incentives55
- Confidence40
build1 publisherOne report A preprint from ServiceNow AI Research and university co-authors reports that a plain web agent with a longer horizon matches or beats AWM, ASI and ReasoningBank, often on fewer tokens.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives45
- Confidence42