OpenAI says better caching and inference let it cut API prices for Sol and Luna by half, and the cost advantage it claims for the cheap tier over the old top tier comes in at one tenth on the benchmark it published and one hundredth in its summary.
Perspective Coverage
8 publishers
- Builder
- Builder 36%
- Operator
- Operator 42%
- Investor
- Investor 22%
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+35
- Incentives70
- Confidence60
OpenAI's own benchmarks put GPT-6 Sol and Luna at half the list price and a fraction of a rival's cost per finished task. For defenders, the thing getting cheaper is autonomous tool calls into business systems.
Reality
- Evidence32
- Adoption38
- Hype gap+34
- Incentives80
- Confidence52
Sol now bills $2 and $10 per million tokens and Luna $0.10 and $0.50, while OpenAI quotes its own benchmark results per task, where the cheap model lands 2.2 points behind Sol on the software engineering test.
Reality
- Evidence42
- Adoption55
- Hype gap+25
- Incentives78
- Confidence52
Sol lists at $2 and $10 per million tokens and Luna at $0.10 and $0.50. The per-task savings OpenAI published come mostly from the lower price, and the cheaper tier scores below its predecessor on computer use.
Publishers:forkast.news · openai.com Reality
- Evidence55
- Adoption32
- Hype gap+30
- Incentives82
- Confidence62
OpenAI has put its paused computer-use model into a few customers' hands, where the speed gain arrives alongside a reasoning trail outside investigators say is harder to follow. The containment work now sits with the customer.
Reality
- Evidence34
- Adoption22
- Hype gap+38
- Incentives70
- Confidence38
DeepSeek V4 Flash Vision Exp undercuts Gemini 3.7 Flash by 3.4x on input tokens and 5.7x on output, then spent 3,467 completion tokens and 30.5 seconds on an invoice both models audited correctly.
Reality
- Evidence56
- Adoption37
- Hype gap+26
- Incentives68
- Confidence55
A published NVFP4 and speculative-decoding config turns a 27B open-weights model into something you can try to serve. The 206.1 tokens per second figure is single-stream and unreplicated.
Reality
- Evidence52
- Adoption28
- Hype gap+14
- Incentives70
- Confidence46
A viral X post said an inference-time text layer put DeepSeek V4 Pro ahead of Fable 5 on every task. The report it points to shows single runs, nine benchmarks, and two losses.
Reality
- Evidence40
- Adoption18
- Hype gap+55
- Incentives62
- Confidence45
Z.ai's August 14 post claims post-training gains for coding agents, but the company's release notes still stop at GLM-5.1 and there is no API endpoint, model identifier or weight download.
Reality
- Evidence42
- Adoption18
- Hype gap+38
- Incentives68
- Confidence46
Z.ai says all of GLM-5.3's coding gains came from post-training on tenfold more long-horizon task environments. The uneven benchmark jumps tell you where that money actually landed.
Reality
- Evidence48
- Adoption30
- Hype gap+24
- Incentives76
- Confidence55