buildConfirmed8 publishers Google's new workhorse Flash is a model-string swap for anyone on AI Gateway, discounted through 31 December 2026. The claim worth testing is reduced tool-calling loop failures, not benchmark deltas.
Perspective Coverage
8 publishers
- Builder
- Builder 54%
- Operator
- Operator 24%
- Investor
- Investor 22%
Reality
- Evidence55
- Adoption45
- Hype gap+25
- Incentives70
- Confidence60
buildConfirmed2 publishers Cognition says its SWE-2 model produced up to 4.8 times more total token throughput on Nvidia's Vera Rubin NVL72 than on GB200, in a test on CoreWeave. The company-run result points to more agent capacity per system, but whether long coding tasks get cheaper depends on what the new systems cost per hour.
Reality
- Evidence40
- Adoption30
- Hype gap+35
- Incentives85
- Confidence60
OpenAI's GPT-6.1 Sol halves cached-input pricing to $0.10 per million tokens and leaves standard rates at $2 and $10. Agents that resend long context collect the saving, while other buyers weigh gains shown mostly in OpenAI's own tests.
Reality
- Evidence55
- Adoption30
- Hype gap+20
- Incentives65
- Confidence55
buildOne report1 publisher Claude Opus 5.5 matched Fable 5.1 on every hidden test in two New Stack coding trials, at $0.75 and $1.42 per run against $1.50 and $1.96. It needed more tokens and more minutes to get there, so the saving holds only for tasks that resemble these.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+15
- Incentives40
- Confidence55
buildOne report1 publisher The Wall Street Journal says Google engineers picked the unreleased Flash model over an Anthropic Opus inside Jetski, a result with real budget implications for agent fleets and no published prompts, judges or Opus version.
Reality
- Evidence38
- Adoption15
- Hype gap+40
- Incentives60
- Confidence45
OpenAI's own benchmarks put GPT-6 Sol and Luna at half the list price and a fraction of a rival's cost per finished task. For defenders, the thing getting cheaper is autonomous tool calls into business systems.
Reality
- Evidence32
- Adoption38
- Hype gap+34
- Incentives80
- Confidence52
Sol lists at $2 and $10 per million tokens and Luna at $0.10 and $0.50. The per-task savings OpenAI published come mostly from the lower price, and the cheaper tier scores below its predecessor on computer use.
Publishers:forkast.news · openai.com Reality
- Evidence55
- Adoption32
- Hype gap+30
- Incentives82
- Confidence62
buildOne report1 publisher The model averages 53 steps a run against SWE-1.7's 127, and Cognition says its mean rollout cost is 64% below Fable 5.1's on a leaderboard Cognition built, runs and grades. Reproducing that takes Devin's harness.
Reality
- Evidence30
- Adoption25
- Hype gap+30
- Incentives80
- Confidence45
Claude Opus 5 is pitched as near-frontier at half the cost. The cost multiple moves with the workload, and every figure on offer is the vendor's own.
Reality
- Evidence30
- Adoption32
- Hype gap+34
- Incentives88
- Confidence52