buildOne report1 publisher Berkeley statisticians Nguyen and Fithian refit METR's time-horizon data and found AI task difficulty barely changes between 2 and 30 minutes of human time. Equal 10x steps on the plot therefore mean unequal capability gains, depending on where on the axis they fall.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap0
- Incentives
- Insufficient
- Confidence40
buildOne report1 publisher Cantina released apex-flash-1, an open-weights vulnerability-research model it says solved 40 of 60 tasks for $2.38, against $74.68 for Claude Opus 5 High. There is no hosted endpoint, so teams download the 321-billion-parameter weights, pay for their own inference and verify the numbers themselves.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+35
- Incentives70
- Confidence40
buildOne report1 publisher Anthropic says its claude-api eval workflow, run by Claude Code, lifted the company's own API skill from about 66% to 88%. That gain is self-reported and holds for another team only if its eval is as realistic and as stable as the loop demands.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+15
- Incentives65
- Confidence45
buildConfirmed2 publishers Broken answer keys and graders that punish correct tool calls drove the verdicts. Epoch AI says it stops each review once it has enough evidence, so the published defect counts are floors.
Reality
- Evidence65
- Adoption
- Insufficient
- Hype gap+5
- Incentives35
- Confidence62
Jev returns typed answers with probability scores and no free-form text, and TypeSafe prices a decision at four hundredths of a cent. Its accuracy figures are scored against other models' probabilities.
Publishers:arize.com
Reality
- Evidence34
- Adoption22
- Hype gap+30
- Incentives68
- Confidence44
buildOne report1 publisher An arXiv paper argues that a hosted model name only routes a request, so a safety finding filed against that name loses its subject as soon as the weights, prompts, classifiers or serving stack change under it.
Reality
- Evidence46
- Adoption12
- Hype gap+25
- Incentives
- Insufficient
- Confidence42
Sentient Labs put a coach model in charge of improving a worker model on a spreadsheet benchmark, and the resulting score jump sat on top of a grading harness that leaked answers in one direction and mismarked correct work in the other.
Reality
- Evidence45
- Adoption12
- Hype gap+25
- Incentives60
- Confidence50
buildOne report1 publisher A sanitized workbook keeps the blank rows, duplicate IDs and text-formatted numbers, and the ten questions put to the assistant all have answers someone already computed by hand. Every formula and cell range gets saved.
Reality
- Evidence32
- Adoption
- Insufficient
- Hype gap+8
- Incentives35
- Confidence45
buildOne report1 publisher Before judging agent-written code, a dev.to author seeded a small C backup client with eight known behaviors and observed it with system-call traces, a loopback capture and a digest check to see which observer caught what.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap−12
- Incentives20
- Confidence55
Alexandr Wang told a Y Combinator audience that a swarm of Meta agents scheduled by cron beat 100 engineers on specific tasks, and he credited the evaluation system for the result. Meta paid $14.3 billion for Scale AI.
Reality
- Evidence20
- Adoption15
- Hype gap+55
- Incentives78
- Confidence60
buildOne report1 publisher ByteDance's Seed team and collaborators had 18 models write their own agent scaffolding and then rewrite it from task feedback. The held-out gains they report land inside the fluctuation band of their own evaluation.
Reality
- Evidence34
- Adoption14
- Hype gap+44
- Incentives61
- Confidence41
buildOne report1 publisher The engineering team's own figures put one pipeline above 90% on clean text retrieval and near 46% on customer-representative files, with the diagnostic trail running back to parsing and chunking rather than to the model.
Reality
- Evidence38
- Adoption35
- Hype gap+22
- Incentives70
- Confidence42
buildOne report1 publisher A $2B Series B rests on synthetic respondents agreeing with human focus groups 85 to 99 percent of the time. The band's floor and ceiling are not the same product.
Reality
- Evidence24
- Adoption28
- Hype gap+46
- Incentives78
- Confidence52