Anthropic and OpenAI are turning to a handful of small, mostly nonprofit evaluators such as METR, Apollo Research and Transluce to check their models. With no federal push for rules, who pays them, what they can see and whom they report to are still open.
Reality
- Evidence55
- Adoption25
- Hype gap+25
- Incentives65
- Confidence50
buildConfirmed22 publishers Anthropic launched Claude Haiku 5.5 at an average price about 75% below Haiku 4.5. The saving varies widely with prompt length, so teams moving classification, support or query traffic need to price their own requests before they switch models.
Perspective Coverage
22 publishers
- Builder
- Builder 47%
- Operator
- Operator 36%
- Investor
- Investor 17%
Reality
- Evidence66
- Adoption30
- Hype gap+25
- Incentives70
- Confidence65
buildConfirmed10 publishers Google priced Gemini 4 Argon at $2 and $10 per million input and output tokens, then released it first to trusted cyber defenders in its Fairwind Program. Teams can budget against those rates now but cannot yet measure the token counts they multiply.
Perspective Coverage
10 publishers
- Builder
- Builder 43%
- Operator
- Operator 29%
- Investor
- Investor 28%
Reality
- Evidence62
- Adoption18
- Hype gap+30
- Incentives68
- Confidence66
buildOne report1 publisher Google prices Gemini 3.8 Flash at $0.75/$3.75 per million input/output tokens through 2026, three-eighths of what partner-only Gemini 4 Argon costs. Building on Flash now works if the later Argon swap moves the thinking settings along with the model name.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap0
- Incentives40
- Confidence60
buildOne report1 publisher Claude Fable 5.1 took an opponent's chess engine in 3 of 10 honeypot games, the same week it solved a 1653 cipher in 44 minutes. Both runs argue for harnesses that enforce tool limits in the sandbox and grade the tool-call trace along with the result.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence50
The top four slots on Artificial Analysis are the advertisement. The line item Anthropic actually moved is the one that scales with how long an agent runs, and its own savings estimate backs out that share at about 60 percent.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+15
- Incentives62
- Confidence60
buildOne report1 publisher Ten Claude Opus 5.5 agents produced the algorithm and a machine-checked proof of its runtime bound in about 15 hours. The theorem covers sparse directed graphs, and the margin over Dijkstra grows as the twelfth root of log n.
Reality
- Evidence52
- Adoption10
- Hype gap+20
- Incentives72
- Confidence55
buildOne report1 publisher The number came out of one harness on OpenRouter at temperature 0.9 with a 64,000-token output cap, and it holds for a bounded translation task while the same model sits mid-table on open-ended terminal work.
Reality
- Evidence57
- Adoption28
- Hype gap+19
- Incentives62
- Confidence60
buildConfirmed2 publishers Vals AI's 141-hour Minecraft run ended with OpenAI's newest model farming potatoes for hours after losing its end-game loot and its spawn point in one explosion, and the recovery policy it came away with was a note to self.
Reality
- Evidence45
- Adoption30
- Hype gap+32
- Incentives68
- Confidence55
Two closes three days apart put $2.85bn of fresh commitments behind the AI stack. The question for anyone raising is which of a16z's two vehicles turns up, because their mandates now overlap at the growth stage.
Reality
- Evidence34
- Adoption55
- Hype gap+15
- Incentives78
- Confidence42
Muse Spark 1.2 costs roughly 18 times less if Meta may train on your prompts. The deepest cut, 75x, sits on cached input, which is where an agent holding your codebase in context spends.
Publishers:deeplearning.ai
Reality
- Evidence56
- Adoption16
- Hype gap+22
- Incentives78
- Confidence48