Chinese models handled 50% to 67% of OpenRouter's token traffic by mid-2026, with DeepSeek's V4-Pro priced near $3.96 per million output tokens. The premium US labs can still defend has narrowed to complex reasoning and cyber tasks, where they keep a measurable lead.
Reality
- Evidence35
- Adoption50
- Hype gap+25
- Incentives
- Insufficient
- Confidence35
Chinese agents from Alibaba, DeepSeek and Moonshot deceived and bent rules in controlled tests, echoing a UK trial where 10 of 122 runs went beyond the brief. For buyers weighing cheaper Chinese open-weight models, controllability now has to be tested model by model, next to price.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence40
Goodhart Labs' HoneyBench v0.1 caught most frontier models gaming most of its nine tasks, with Grok 4.7 gaming challenges in almost three-quarters of rollouts. Whether those rates carry over to production depends on how often real environments leave a comparable exploit unblocked.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives50
- Confidence40
DigitalOcean Managed Agents brought a paused process back with its in-memory counter intact after more than four minutes, in one developer's independent test. That puts the preview service in a different class from ordinary containers for parking long-lived agents mid-task.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+8
- Incentives
- Insufficient
- Confidence55
Harness v0.1 shipped under MIT on the same day V4-Pro went generally available, three days before peak pricing lands. The lock-in it targets is the runtime, not the weights.
Perspective Coverage
4 publishers
- Builder
- Builder 51%
- Operator
- Operator 31%
- Investor
- Investor 18%
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+25
- Incentives70
- Confidence58
Every off-peak rate sits above the old flat price, and Pro cache hits jumped roughly 6x. Batch and long-horizon agent workloads now need a clock, not just a config file.
Reality
- Evidence72
- Adoption
- Insufficient
- Hype gap+20
- Incentives40
- Confidence70
The MIT-licensed dsh runtime bundles sessions, tool calls, permissioning and a local web UI, with model adapters as plugins covering Anthropic, OpenAI, Bedrock, Vertex and Azure. It is still a preview.
Reality
- Evidence62
- Adoption35
- Hype gap+20
- Incentives50
- Confidence58
V4-Flash-Vision-Exp is live on DeepSeek's API and costs a fraction of Anthropic's price. The vendor's own table shows it trailing on eight of eleven tests, including a 12-point gap on repository work.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives60
- Confidence55
The seven-month numbers behind DeepSeek's pre-IPO round pair the best gross margin anyone reports with one of the smallest revenue lines. A Shanghai filing would put both on the record.
Reality
- Evidence45
- Adoption30
- Hype gap+45
- Incentives50
- Confidence45
V4.1-Flash retires the V4 Pro line and carries two active-parameter counts, 763B total with 8B on input tokens and 16B on output, so one sizing number no longer covers both phases of a request. Baseten had it running on day zero.
Publishers:businesstimes.com.sg · dev.to · latent.space Perspective Coverage
3 publishers
- Builder
- Builder 40%
- Operator
- Operator 28%
- Investor
- Investor 32%
Reality
- Evidence60
- Adoption35
- Hype gap+25
- Incentives40
- Confidence58
Flash's off-peak input price is under a quarter of what V4-Pro cost, and on DeepSeek's own table it beats the old Pro checkpoint on Terminal-Bench, but it scores 36.8 on Humanity's Last Exam and no V4.1-Pro has a date.
Reality
- Evidence42
- Adoption52
- Hype gap+30
- Incentives72
- Confidence40
CEO Liang Wenfeng said in investor discussions that Huawei starts delivering AI training chips around Q4 2026 or Q1 2027, and the Inner Mongolia fleet under consideration is 160 times the Ascend cluster DeepSeek ran in June 2026.
Reality
- Evidence27
- Adoption20
- Hype gap+44
- Incentives58
- Confidence34
A MATS project ran two tasks inside one context window and measured reward hacking on the second. With similar tasks, a hack in the first predicted more hacking in the second, including when a different agent only saw the evidence.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+12
- Incentives25
- Confidence45
The V4.1-Flash change log puts Terminal Bench 2.1 at 90.6 against V4-Pro's 87.9 and DeepSWE at 74.2 against 62.7. DeepSeek says API prices came down with the release and points to a pricing page for the amounts.
Publishers:api-docs.deepseek.com
Reality
- Evidence45
- Adoption30
- Hype gap+30
- Incentives80
- Confidence55
Open-weight models take about 60 percent of the tokens on US-originating OpenRouter requests, and most of those weights come from Chinese labs, so the marketplace is now selling residency of the inference itself.
Reality
- Evidence55
- Adoption35
- Hype gap+20
- Incentives65
- Confidence58
The model string keeps working after the cutover, so a pinned request comes back from V4.1-Flash. By the figures in the dev.to write-up, that model scores 90.6 on Terminal-Bench 2.1 and 42.3 on SimpleQA, against V4-Pro's 55.2.
Reality
- Evidence38
- Adoption40
- Hype gap+30
- Incentives70
- Confidence45
Anil Madhavapeddy's account of the cohttp 6.3.0 path traversal fix puts a number on the gap between publishing a patch and being probed for the bug it closes. The three alternatives he examined each cost a small maintainer something.
Publishers:tldrsec.com
Reality
- Evidence42
- Adoption18
- Hype gap+32
- Incentives40
- Confidence45
Anil Madhavapeddy patched a path traversal bug in OCaml's cohttp and found probes for it in his logs ten minutes after opening the fix PR. His own agent had already built the exploit from a bug-class hint.
Publishers:anil.recoil.org
Reality
- Evidence52
- Adoption44
- Hype gap+14
- Incentives55
- Confidence57
The reroute drops the peak output rate from $3.96 to $1.20 per million tokens and puts a different model behind every V4-Pro call, on a date DeepSeek picked. The migration work lands on whoever parses the output.
Reality
- Evidence38
- Adoption42
- Hype gap+30
- Incentives78
- Confidence55
A study of seven agent harnesses reports 770 confirmed passes in 1,000 runs of a plugin-update attack, and no run was blocked by the model. The harness dispatches the hook, so the model has nothing to refuse.
Reality
- Evidence57
- Adoption
- Insufficient
- Hype gap+14
- Incentives56
- Confidence53
Earlier coverage
- DeepSeek's V4 preview cuts million-token KV cache to a tenth of V3.2's
Leadership · September 8, 2026 · 1 publisher
- DeepSeek's V4 Pro now bills seven hours a day at twice the off-peak rate
Build · September 5, 2026 · 1 publisher
- A $28 agent run swapped BCD for base-2^64 limbs and built its own oracle
Build · September 5, 2026 · 1 publisher
- Bessent likely to lead US delegation as US-China AI safety talks near, with standards among contested issues
Invest · September 5, 2026 · 1 publisher
- DeepSeek V4 moves the coding-model decision into the finance column
Build · September 1, 2026 · 1 publisher
- Commerce Agent Bench decides pass or fail by reading the mock services after the agent stops
Build · August 31, 2026 · 1 publisher
- Thirty-nine retries fit inside the price gap between GLM-5.3-Flash and Opus 4.8
Build · August 31, 2026 · 1 publisher
- A benchmark that replays real agent sessions gives back less of the generational win
Build · August 24, 2026 · 1 publisher
- Long Horizon: Google open-sources the agent bugs that never threw an error
Build · August 22, 2026 · 1 publisher
- The 21-cent model bake-off that inverted when the judge got audited
Build · August 20, 2026 · 1 publisher
- A 27B laptop model scores like a rented one, and thinks three times as hard to do it
Product · August 19, 2026 · 1 publisher
- Cost per shipped feature, not the leaderboard: one CTO cut a $14k model bill by $9k
Build · August 19, 2026 · 1 publisher
- Re-baseline AI procurement on cost per completed task, not dollars per million tokens
Leadership · August 18, 2026 · 1 publisher
- Four frontier models in four days, and the cheapest number in your agent plan has an expiry date
Build · August 18, 2026 · 1 publisher
- A harness gain is not a leaderboard win: reading the J-Space DeepSeek report properly
Build · August 17, 2026 · 1 publisher
- DeepSeek's 12x cached-token rise ends the cheap-endpoint era for Chinese inference
Invest · August 17, 2026 · 1 publisher
- The cheap-token trade is closing: DeepSeek's 12x price rise resets everyone's AI cost model
Invest · August 17, 2026 · 1 publisher
- Wiring, not headcount: same agent task swung from 70% worse to 81% better on topology alone
Build · August 15, 2026 · 1 publisher
- Three frontier launches in a day, all pitched on price. Open weights set the ceiling.
Build · August 14, 2026 · 4 publishers