The replay only works on runs Warp's infrastructure already recorded, and correctness is graded by a judge model against a rubric you write. Warp's own 30-task bake-off cost $2,130.57. It finished in under four hours.
Publishers:runtimewire.com · warp.dev Reality
- Evidence40
- Adoption15
- Hype gap+35
- Incentives80
- Confidence55
Dioxus engineers already had code running inside Devin's command line, and they arrive with a cloud harness that snapshots and forks virtual machines on five operating systems. Neither side called it an acquisition.
Reality
- Evidence55
- Adoption40
- Hype gap+15
- Incentives60
- Confidence55
Berkeley's RDI center attacked the step where each benchmark computes its score, and without solving a task its own scorecard reports 100% on five of the eight, about 98% on GAIA and 73% on OSWorld.
Publishers:rdi.berkeley.edu
Reality
- Evidence45
- Adoption35
- Hype gap+20
- Incentives55
- Confidence50
A 300-trial study swapped Goose, OpenCode and OpenHands-SDK under Qwen 3.6 Plus and MiniMax M2.5, and reports that the scaffold sets tokens per solved task and the failure mode while the score barely moves.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+20
- Incentives30
- Confidence58
The company benchmarked coding agents on real tasks against its own multi-million line codebase and found that per-token price predicted almost nothing about what a finished task cost. GLM 5.2 came in at $1.28.
Reality
- Evidence58
- Adoption38
- Hype gap+20
- Incentives70
- Confidence55
The Kotlin Benchmark grades agents on 105 verified repository tasks. Its token column shows setups that solve within a few tasks of each other burning between 66,000 and 777,000 tokens per fix.
Reality
- Evidence62
- Adoption30
- Hype gap+22
- Incentives60
- Confidence58
The efficiency claims are self-reported and measured against DeepSeek's own prior model, but the weights are MIT-licensed and already pulled 1.78 million times, which is what turns a ratio into a number a buyer can carry into a renewal.
Reality
- Evidence45
- Adoption55
- Hype gap+35
- Incentives80
- Confidence58
The study counts only complexity and dead code, because those are the two pyscn metrics that stay exact when you analyze just the files a patch touched. The human's own commit trips the same rule 24% of the time.
Reality
- Evidence55
- Adoption20
- Hype gap+12
- Incentives62
- Confidence60
SWE-Gate scores agent patches at two gates instead of one. The third that clear tests but fail mined review rules are measured against a pipeline that runs tests and nothing else, while most teams run a pipeline with additional checks beyond that.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+45
- Incentives25
- Confidence35
A dev.to writeup attributes 612.9M tokens and $28.35 to four days of DeepSeek Harness work on a C# arbitrary-precision library, and the part worth reading is how the agent checked its own rewrite.
Reality
- Evidence35
- Adoption15
- Hype gap+25
- Incentives55
- Confidence40
Six models from 0.9B to 375B, four of them pretrained on one shared token sequence, mean IFM's sparse-versus-dense efficiency claim can be tested by outsiders at a scale they can actually afford to re-run.
Reality
- Evidence44
- Adoption18
- Hype gap+16
- Incentives68
- Confidence52
A dev.to writeup grades a bug-fixing agent on its tool calls rather than its patch. Thirty documented runs cost roughly a dollar; the expensive input is the hand-built dataset behind them.
Reality
- Evidence34
- Adoption11
- Hype gap−8
- Incentives24
- Confidence41
A practitioner's field notes put curated agent evaluations on one axis and production intent resolution on another. The cost of confusing them gets billed per session, not per query.
Reality
- Evidence18
- Adoption
- Insufficient
- Hype gap+22
- Incentives42
- Confidence30
PlannerCritic's author tried to inject his own engine. The blocks arrived as feasibility verdicts rather than safety strings, which is a design worth copying and a limit worth reading closely.
Reality
- Evidence44
- Adoption12
- Hype gap+24
- Incentives72
- Confidence41
PlannerCritic's sweeps are cheap enough to gate every release. The harder problem is telling a clean run from a rig that quietly counted nothing.
Reality
- Evidence34
- Adoption11
- Hype gap+24
- Incentives68
- Confidence33
A new diagnostic benchmark treats the execution layer as something to vary rather than a fixed backdrop. Its conclusion is that agent capability belongs to a model-harness pair.
Reality
- Evidence42
- Adoption10
- Hype gap+22
- Incentives55
- Confidence40
FreeToken's preprint puts DeepSeek-V4-Flash on a single RTX 5090 desktop. The GPU holds about a sixth of the machine's memory, and nobody outside the team has run the benchmark yet.
Reality
- Evidence44
- Adoption12
- Hype gap+16
- Incentives58
- Confidence52
Veracode ran more than 150 models over 80 tasks and found 45% of the output carries a known weakness. The level is bad; the flat two-year trend is the part that changes your plan.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+25
- Incentives78
- Confidence48
ProdCodeBench builds tasks from real assistant sessions in an industrial monorepo. Its authors report that models using validation tools more heavily solve more, which argues for scoring tool discipline.
Reality
- Evidence44
- Adoption21
- Hype gap+16
- Incentives58
- Confidence42
Optima lets buyers build benchmarks from their own datasets and agent traces, then scores candidate models on quality, cost per task and time per task.
Reality
- Evidence34
- Adoption16
- Hype gap+22
- Incentives71
- Confidence33