build11 publishersConfirmed Anthropic launched Claude Haiku 5.5 at an average price about 75% below Haiku 4.5. The saving varies widely with prompt length, so teams moving classification, support or query traffic need to price their own requests before they switch models.
Perspective Coverage
8 publishers
- Builder
- Builder 53%
- Operator
- Operator 28%
- Investor
- Investor 19%
Reality
- Evidence68
- Adoption30
- Hype gap+20
- Incentives65
- Confidence70
build1 publisherOne report H Company's open-weight Holo4 27B handles GUIs, code, MCP servers and REST APIs in one model, scoring 61.7% on OSWorld 2.0 at an estimated $1.22 per task. The choice between specialists happens during training, so agent teams could drop the router from their stack if the scores hold on their own workloads.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+25
- Incentives
- Insufficient
- Confidence40
OpenAI lists GPT-6.1 Sol at $2 per million input tokens and $10 per million output; the 10-cent figure in early coverage is its cached-input rate. Teams moving work off Astra should budget on the list rates and OpenAI's per-task costs.
Perspective Coverage
5 publishers
- Builder
- Builder 44%
- Operator
- Operator 34%
- Investor
- Investor 22%
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+35
- Incentives70
- Confidence60
OpenAI's GPT-6.1 Sol halves cached-input pricing to $0.10 per million tokens and leaves standard rates at $2 and $10. Agents that resend long context collect the saving, while other buyers weigh gains shown mostly in OpenAI's own tests.
Reality
- Evidence55
- Adoption30
- Hype gap+20
- Incentives65
- Confidence55
build8 publishersConfirmed Google's third Flash release in six weeks keeps the $0.75/$3.75 rate card. But the model also spends more tokens per task. Both numbers in your cost model are moving before the price even changes.
Perspective Coverage
8 publishers
- Builder
- Builder 53%
- Operator
- Operator 29%
- Investor
- Investor 18%
Reality
- Evidence58
- Adoption35
- Hype gap+22
- Incentives72
- Confidence62
build1 publisherOne report OpenAI's spec lets gpt-6-astra take 922,000 input tokens, but requests above 272,000 move to higher long-context rates. Provisioning a model that drives a desktop, a shell and MCP servers starts with that threshold.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives55
- Confidence40
OpenAI says better caching and inference let it cut API prices for Sol and Luna by half, and the cost advantage it claims for the cheap tier over the old top tier comes in at one tenth on the benchmark it published and one hundredth in its summary.
Perspective Coverage
8 publishers
- Builder
- Builder 36%
- Operator
- Operator 42%
- Investor
- Investor 22%
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+35
- Incentives70
- Confidence60
OpenAI's own benchmarks put GPT-6 Sol and Luna at half the list price and a fraction of a rival's cost per finished task. For defenders, the thing getting cheaper is autonomous tool calls into business systems.
Reality
- Evidence32
- Adoption38
- Hype gap+34
- Incentives80
- Confidence52
Sol now bills $2 and $10 per million tokens and Luna $0.10 and $0.50, while OpenAI quotes its own benchmark results per task, where the cheap model lands 2.2 points behind Sol on the software engineering test.
Reality
- Evidence42
- Adoption55
- Hype gap+25
- Incentives78
- Confidence52
Sol lists at $2 and $10 per million tokens and Luna at $0.10 and $0.50. The per-task savings OpenAI published come mostly from the lower price, and the cheaper tier scores below its predecessor on computer use.
Publishers:forkast.news · openai.com Reality
- Evidence55
- Adoption32
- Hype gap+30
- Incentives82
- Confidence62
Berkeley's RDI center attacked the step where each benchmark computes its score, and without solving a task its own scorecard reports 100% on five of the eight, about 98% on GAIA and 73% on OSWorld.
Publishers:rdi.berkeley.edu
Reality
- Evidence45
- Adoption35
- Hype gap+20
- Incentives55
- Confidence50
build1 publisherOne report Sonnet 4.5 still leads GPT-5 on the coding leaderboards, and GPT-5 lists about 46 percent below it on a 5:1 token mix. Anthropic's current Sonnet undercuts both of Sonnet 4.5's list prices, and that complicates a routing plan built on the older pair.
Reality
- Evidence40
- Adoption20
- Hype gap+15
- Incentives55
- Confidence45
build1 publisherOne report Mininglamp published NavEval scores for its own model on its own benchmark. Across the three entries, the spread tracks how each stack reads a page. It is not a case of specialists beating frontier models.
Reality
- Evidence24
- Adoption12
- Hype gap+46
- Incentives86
- Confidence58
build1 publisherOne report A new diagnostic benchmark treats the execution layer as something to vary rather than a fixed backdrop. Its conclusion is that agent capability belongs to a model-harness pair.
Reality
- Evidence42
- Adoption10
- Hype gap+22
- Incentives55
- Confidence40