Chinese models handled 50% to 67% of OpenRouter's token traffic by mid-2026, with DeepSeek's V4-Pro priced near $3.96 per million output tokens. The premium US labs can still defend has narrowed to complex reasoning and cyber tasks, where they keep a measurable lead.
Reality
- Evidence35
- Adoption50
- Hype gap+25
- Incentives
- Insufficient
- Confidence35
Six LLMs on a 12-task Kaggle security benchmark all caught SQL injection, hardcoded keys and pickle RCE, but Gemini 3.7 Flash missed a path traversal. With one scenario per flaw class, the run shows which textbook patterns the models know and says little about trusting one to review real code.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+40
- Incentives40
- Confidence35
Grok 4.20 followed an injected wrong step about five times as often with reasoning mode on, in a 15-question Kaggle benchmark entry. On the Anthropic test it uses, that makes the trace a better record of how the model answered and a weaker guard against a bad step.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+25
- Incentives20
- Confidence35
Blog vs Bytecode, a 28-item Kaggle benchmark, graded empty proxy responses as wrong and scored DeepSeek-R1 at 17% until a second gateway showed 100%. Once capture was fixed, frontier models lost points by flagging sound code, while a small Gemma model missed most of the planted flaws.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+10
- Incentives30
- Confidence45
Jalapeno's lead is measured against last generation and the volumes are tiny, but a first-pass ASIC clearing Nvidia, AMD and Google parts reprices the design barrier, not the supply chain.
Perspective Coverage
7 publishers
- Builder
- Builder 30%
- Operator
- Operator 24%
- Investor
- Investor 46%
Reality
- Evidence50
- Adoption8
- Hype gap+40
- Incentives65
- Confidence55
Reading a 600 GB checkpoint from local NVMe at 7 GB/s takes about 86 seconds, so the network stops gating scale-out on warm nodes. The first fill still costs the full download. The slowest node sets it.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+35
- Incentives85
- Confidence55
NVIDIA measured TensorRT LLM holding 96.1 to 98.2 percent of its non-confidential output throughput on Blackwell, and it got there by unpinning host memory on the affected paths, moving decode readback off the scheduler thread, and timing kernel tactics with the GPU's global timer instead of CUDA events.
Reality
- Evidence58
- Adoption22
- Hype gap+10
- Incentives80
- Confidence55
A dev.to walkthrough of DeepSeek's GRPO puts a 70B PPO training loop over 600GB of VRAM before any sharding. The critic-free version deletes the most expensive network and pays for it in sampled completions.
Reality
- Evidence32
- Adoption
- Insufficient
- Hype gap+30
- Incentives42
- Confidence45
METR also had the agent match the median human expert on five AI R&D tasks, using 32 hours of wall clock against the humans' eight, assembled from attempts of two hours or less. The confidence intervals still overlap the other public models it has tested.
Reality
- Evidence58
- Adoption25
- Hype gap−10
- Incentives40
- Confidence55
MLPerf Inference v6.1 preview submissions put Vera Rubin NVL72 at up to 3.7x GB300 on Qwen3-VL and up to 2.5x on DeepSeek-R1, on two different inference frameworks. The four-rack 99% scaling result is an offline number.
Reality
- Evidence45
- Adoption30
- Hype gap+35
- Incentives85
- Confidence55
Anthropic tested whether a model's visible reasoning names what actually drove its answer, and for two reasoning models it usually did not. Teams treating traces as an audit record are relying on that property.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap−15
- Incentives62
- Confidence60
DeepSeek's first CFO arrives with CITIC Securities engaged and a 2027 listing as the goal. The near-term cash is a 50 billion yuan private round, and the two published accounts of June's valuation differ by nearly seven times.
Reality
- Evidence42
- Adoption33
- Hype gap+38
- Incentives66
- Confidence48
The 1.7x to 3.6x latency range is set by the baseline systems, not the chip, and the report's own publication date is unsettled. Read it as direction, not evidence.
Perspective Coverage
9 publishers
- Builder
- Builder 41%
- Operator
- Operator 31%
- Investor
- Investor 28%
Reality
- Evidence52
- Adoption14
- Hype gap+38
- Incentives82
- Confidence68
American agencies say six Chinese labs bought bulk subscriptions to US models and trained on the outputs since 2024. Enterprise buyers are meanwhile paying a fifth as much for models that clear most of their engineering work.
Reality
- Evidence45
- Adoption55
- Hype gap+25
- Incentives70
- Confidence45
A preprint holds 1,520 benchmark responses constant and varies one sentence about what a low score will do to the model being scored. The judges get more lenient, and their reasoning traces never mention the sentence.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+18
- Incentives45
- Confidence52
CISA, the NSA and the FBI want US model providers to catch industrial-scale distillation hiding inside traffic that looks like a busy enterprise customer, then quietly degrade the answers, which makes false positives a product decision.
Reality
- Evidence46
- Adoption
- Insufficient
- Hype gap+32
- Incentives70
- Confidence42
Sarah Friar told Goldman Sachs' technology conference that a pricier model can be cheaper when it needs fewer tries, a claim Artificial Analysis puts at an eleven-to-one price spread against seven index points running the other way.
Reality
- Evidence34
- Adoption38
- Hype gap+30
- Incentives80
- Confidence45
The joint NSA, CISA and FBI advisory names six China-based companies and describes billions of tokens taken through paid APIs, cloud resellers and proxy transfer stations, which turns billing telemetry into a detection surface.
Reality
- Evidence48
- Adoption30
- Hype gap+32
- Incentives76
- Confidence55
Mithil Vakde's from-scratch transformer landed one point behind TRM on the public eval. His own ablations put 20 of those 44 points on two representation choices rather than on any amount of compute.
Reality
- Evidence44
- Adoption18
- Hype gap+16
- Incentives72
- Confidence41
A team adapting the implicit association test to reasoning traces found four of five models working harder on association-incompatible prompts, which puts a measurable bias signal in the process rather than only in the answer.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+14
- Incentives45
- Confidence55
Earlier coverage
- DeepSeek V4 moves the coding-model decision into the finance column
Build · September 1, 2026 · 1 publisher
- X discloses Chinese bot farm accounts posting on AI data centre and energy debate
Invest · August 30, 2026 · 2 publishers
- Speculative decoding gets scaling laws you can size a draft model against
Build · August 28, 2026 · 1 publisher
- SPEED-Bench re-tests speculative decoding at the batch size you actually serve
Build · August 27, 2026 · 1 publisher
- Nvidia's power pitch: 40,000 Rubin GPUs in 100MW, or 2.5kW a GPU at the meter
Build · August 26, 2026 · 1 publisher
- The judge went synthetic first, which tells you which part of your pipeline is next
Build · August 22, 2026 · 1 publisher
- Cost per shipped feature, not the leaderboard: one CTO cut a $14k model bill by $9k
Build · August 19, 2026 · 1 publisher
- Nature Perspective: patching one fact into a model leaves the reasoning around it broken
Science · August 16, 2026 · 1 publisher