OpenAI has cut input prices on its Luna models from $1.00 to $0.10 per million tokens since July 30, over two rounds of reductions. Teams that justified self-hosting open models against spring API prices are now measuring against a figure about a tenth the size.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+30
- Incentives45
- Confidence40
Harness v0.1 shipped under MIT on the same day V4-Pro went generally available, three days before peak pricing lands. The lock-in it targets is the runtime, not the weights.
Perspective Coverage
4 publishers
- Builder
- Builder 51%
- Operator
- Operator 31%
- Investor
- Investor 18%
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+25
- Incentives70
- Confidence58
Every off-peak rate sits above the old flat price, and Pro cache hits jumped roughly 6x. Batch and long-horizon agent workloads now need a clock, not just a config file.
Reality
- Evidence72
- Adoption
- Insufficient
- Hype gap+20
- Incentives40
- Confidence70
LiveReview's Maneshwar says Gemini File Search cost over 50 cents a review run because a reasoning model performed each search. A local index and a cheaper model brought runs down to 4 cents.
Reality
- Evidence45
- Adoption10
- Hype gap+15
- Incentives35
- Confidence40
V4-Flash-Vision-Exp is live on DeepSeek's API and costs a fraction of Anthropic's price. The vendor's own table shows it trailing on eight of eleven tests, including a 12-point gap on repository work.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives60
- Confidence55
Enterprises buy the cheapest model that clears their bar. On Ramp's July billing data, that leaves Anthropic's flagship with about an eighth of its maker's platform spend.
Reality
- Evidence55
- Adoption25
- Hype gap+30
- Incentives40
- Confidence55
V4.1-Flash retires the V4 Pro line and carries two active-parameter counts, 763B total with 8B on input tokens and 16B on output, so one sizing number no longer covers both phases of a request. Baseten had it running on day zero.
Publishers:businesstimes.com.sg · dev.to · latent.space Perspective Coverage
3 publishers
- Builder
- Builder 40%
- Operator
- Operator 28%
- Investor
- Investor 32%
Reality
- Evidence60
- Adoption35
- Hype gap+25
- Incentives40
- Confidence58
Flash's off-peak input price is under a quarter of what V4-Pro cost, and on DeepSeek's own table it beats the old Pro checkpoint on Terminal-Bench, but it scores 36.8 on Humanity's Last Exam and no V4.1-Pro has a date.
Reality
- Evidence42
- Adoption52
- Hype gap+30
- Incentives72
- Confidence40
GLM-5.3-Flash leads agentic terminal work, DeepSeek V4 Flash is billed as the cheapest per token, and a 2.52B MiniCPM5-2B runs locally under Apache 2.0. The comparison flags most of those numbers as vendor-reported.
Reality
- Evidence34
- Adoption27
- Hype gap+26
- Incentives58
- Confidence41
Sentient Labs put a coach model in charge of improving a worker model on a spreadsheet benchmark, and the resulting score jump sat on top of a grading harness that leaked answers in one direction and mismarked correct work in the other.
Reality
- Evidence45
- Adoption12
- Hype gap+25
- Incentives60
- Confidence50
One endpoint fronts more than 300 models, but the company running the GPUs picks the inference engine and the quantization. A dev.to writeup says the quality gap that follows turns up in the response body, while the status code still reads success.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+25
- Incentives
- Insufficient
- Confidence45
The V4.1-Flash change log puts Terminal Bench 2.1 at 90.6 against V4-Pro's 87.9 and DeepSWE at 74.2 against 62.7. DeepSeek says API prices came down with the release and points to a pricing page for the amounts.
Publishers:api-docs.deepseek.com
Reality
- Evidence45
- Adoption30
- Hype gap+30
- Incentives80
- Confidence55
Shrivu Shankar rented cloud GPUs, pointed about 100 abliterated open-source agents at his own name, and five hours later two weak passwords and three old side projects had fallen while his Gmail, 1Password and bank held.
Publishers:blog.sshh.io
Reality
- Evidence44
- Adoption22
- Hype gap+9
- Incentives36
- Confidence55
The model card for this 552B-parameter Mixture-of-Experts release puts the global KV cache at 890 bytes per token, about a quarter of the previous Flash generation, and every figure in it is DeepSeek's own.
Reality
- Evidence58
- Adoption20
- Hype gap+18
- Incentives82
- Confidence60
The reroute drops the peak output rate from $3.96 to $1.20 per million tokens and puts a different model behind every V4-Pro call, on a date DeepSeek picked. The migration work lands on whoever parses the output.
Reality
- Evidence38
- Adoption42
- Hype gap+30
- Incentives78
- Confidence55
Perplexity pairs about 190 million web pages with 69,721 agent-written queries, which is closer to production than most retrieval tests get, and keeps the corpus, queries and labels private so it stays the only party able to run it.
Reality
- Evidence50
- Adoption15
- Hype gap+15
- Incentives88
- Confidence55
The efficiency claims are self-reported and measured against DeepSeek's own prior model, but the weights are MIT-licensed and already pulled 1.78 million times, which is what turns a ratio into a number a buyer can carry into a renewal.
Reality
- Evidence45
- Adoption55
- Hype gap+35
- Incentives80
- Confidence58
Qwen3.8-Flash-Next puts 36 Gated DeltaNet layers and 12 sparse-attention layers on Hugging Face, which means the retrieval budget Qwen4 will inherit is something you can measure against your own traces now.
Perspective Coverage
5 publishers
- Builder
- Builder 52%
- Operator
- Operator 28%
- Investor
- Investor 20%
Reality
- Evidence58
- Adoption52
- Hype gap+32
- Incentives76
- Confidence71
The old flat rate became a peak rate that blends 3.6x higher, with a half-price window covering seventeen hours, so inference cost modelling now depends on which UTC hour the traffic lands in.
Reality
- Evidence46
- Adoption44
- Hype gap+12
- Incentives45
- Confidence50
Chinese banks and telcos are packaging inference the way airlines package miles, though the analyst closest to the trend calls the consumer bundles a supply-led experiment whose users never see a balance.
Reality
- Evidence54
- Adoption61
- Hype gap+14
- Incentives71
- Confidence57
Earlier coverage
- DeepSeek V4 moves the coding-model decision into the finance column
Build · September 1, 2026 · 1 publisher
- Thirty-nine retries fit inside the price gap between GLM-5.3-Flash and Opus 4.8
Build · August 31, 2026 · 1 publisher
- Peak-hour pricing pushes DeepSeek's new vision model past Gemini on the invoice test
Build · August 31, 2026 · 2 publishers
- Five buyers now owe Nvidia $44bn of the demand its guidance promised
Invest · August 27, 2026 · 1 publisher
- Harness choice moved token use 83-fold with the model held constant
Build · August 27, 2026 · 1 publisher
- Stealth is now a launch strategy: Zhipu's Ox Alpha topped the charts before it had a name
Product · August 26, 2026 · 1 publisher
- Ox Alpha passes the Xinjiang test and fails on Xi: seven topics, 83 points apart
Build · August 25, 2026 · 1 publisher
- DeepSeek V4 doubled its OpenRouter token share, and the bill it displaced was ~130x bigger
Invest · August 25, 2026 · 1 publisher
- A 284B model at 25 tokens a second on one 5090, and 192 GiB of DDR5 doing the quiet part
Build · August 22, 2026 · 1 publisher
- A 284B model in 3.2GB of RAM turns sparse MoE into a disk bandwidth problem
Build · August 21, 2026 · 1 publisher
- Cost per shipped feature, not the leaderboard: one CTO cut a $14k model bill by $9k
Build · August 19, 2026 · 1 publisher
- Re-baseline AI procurement on cost per completed task, not dollars per million tokens
Leadership · August 18, 2026 · 1 publisher
- Your Multi-Key Failover Is The Most Expensive Line On Your Coding Agent Bill
Build · August 18, 2026 · 1 publisher
- A Beijing bar gives away DeepSeek tokens. Your per-token price list is the collateral damage.
Product · August 17, 2026 · 1 publisher
- Four months of A100 bills say self-hosting is a utilization bet, not a cost saving
Build · August 16, 2026 · 1 publisher
- DeepSeek V4 Flash costs a tenth as much and passes 53.8% of agent tasks
Invest · August 16, 2026 · 1 publisher
- The zero-false-accept memory result belonged to the proxy, not the model
Build · August 14, 2026 · 1 publisher