Kaggle benchmark results posted on dev.to report that most of 30 vision models reading 336 synthetic Grafana-style panels found the peak but misread the clock. Copilot incident timelines drafted from screenshots need their start times checked by hand.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives30
- Confidence45
Generating the interface as video instead of rendering it from code is a serious research bet. Runway's own preview still lists legible text and long-session coherence as open problems, which is roughly where ordinary interfaces begin.
Perspective Coverage
5 publishers
- Builder
- Builder 47%
- Operator
- Operator 37%
- Investor
- Investor 16%
Reality
- Evidence45
- Adoption5
- Hype gap+35
- Incentives65
- Confidence60
A MATS project ran two tasks inside one context window and measured reward hacking on the second. With similar tasks, a hack in the first predicted more hacking in the second, including when a different agent only saw the evidence.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+12
- Incentives25
- Confidence45
The team behind a nine-agent code generator reports about 30,000 tokens and 15 to 20 Pro-tier model calls per request, and attributes its reliability to strict output schemas and context the agents cannot decline to read.
Reality
- Evidence40
- Adoption28
- Hype gap+15
- Incentives55
- Confidence42
Mithil Vakde's from-scratch transformer landed one point behind TRM on the public eval. His own ablations put 20 of those 44 points on two representation choices rather than on any amount of compute.
Reality
- Evidence44
- Adoption18
- Hype gap+16
- Incentives72
- Confidence41
The company's own benchmark keeps every flagship model it tested below 60 percent completion across 2140 CRM task instances, and the fix its researchers report is domain knowledge somebody inside your org has to write down.
Reality
- Evidence52
- Adoption22
- Hype gap−10
- Incentives80
- Confidence55
CRMArena-Pro reports about 58% single-turn success and 35% multi-turn over nineteen expert-validated tasks, and the per-skill breakdown inside those averages is the part that should decide where an agent gets pointed.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap−12
- Incentives52
- Confidence58
ArmBench-ASR v0.1 ranks nearly 30 systems on 20.7 hours of Armenian audio. The headline order flips on read speech, and every model degrades badly on movie dialogue.
Reality
- Evidence57
- Adoption22
- Hype gap+12
- Incentives58
- Confidence54
A Secure Code Warrior and RMIT study of six frontier models across 11 frameworks found no universal winner and no link between token cost and secure output.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+20
- Incentives68
- Confidence40
Google Research reports Gemini-3-Pro and GPT-5 encode 95-98% of tested facts yet fail to recall 26-34% of them, moving the fix from pretraining scale toward post-training and inference.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+20
- Incentives68
- Confidence48