NVIDIA measured TensorRT LLM holding 96.1 to 98.2 percent of its non-confidential output throughput on Blackwell, and it got there by unpinning host memory on the affected paths, moving decode readback off the scheduler thread, and timing kernel tactics with the GPU's global timer instead of CUDA events.
Reality
- Evidence58
- Adoption22
- Hype gap+10
- Incentives80
- Confidence55
MLPerf Inference v6.1 preview submissions put Vera Rubin NVL72 at up to 3.7x GB300 on Qwen3-VL and up to 2.5x on DeepSeek-R1, on two different inference frameworks. The four-rack 99% scaling result is an offline number.
Reality
- Evidence45
- Adoption30
- Hype gap+35
- Incentives85
- Confidence55
SemiAnalysis says its AgentX benchmark clocked Nvidia's Vera Rubin NVL72 at up to seven times Blackwell's token throughput per megawatt on pre-release software, and puts the profit gain at over two times per gigawatt.
Reality
- Evidence46
- Adoption29
- Hype gap+16
- Incentives67
- Confidence41
SageMaker Inference now sends requests that begin with the same tokens to the same instance. On AWS's seven-node Llama 3.1 70B benchmark that lifted the KV cache hit rate from roughly 25 percent to 82 percent.
Reality
- Evidence48
- Adoption22
- Hype gap+22
- Incentives85
- Confidence52
Qwen3.8-Flash-Next puts 36 Gated DeltaNet layers and 12 sparse-attention layers on Hugging Face, which means the retrieval budget Qwen4 will inherit is something you can measure against your own traces now.
Perspective Coverage
5 publishers
- Builder
- Builder 52%
- Operator
- Operator 28%
- Investor
- Investor 20%
Reality
- Evidence58
- Adoption52
- Hype gap+32
- Incentives76
- Confidence71
Ten Chinese model variants now answer an OpenAI tool-calling payload well enough to return a valid response, which is the easy part. The streaming deltas and call arity behind it are where an agent migration spends its budget.
Reality
- Evidence36
- Adoption47
- Hype gap+14
- Incentives
- Insufficient
- Confidence43
Speculative decoding speedups depend on the data, and most published ones come from high-level scripts on narrow datasets. SPEED-Bench's authors argue the honest measurement happens inside vLLM or TensorRT-LLM, across concurrencies.
Reality
- Evidence42
- Adoption21
- Hype gap+18
- Incentives58
- Confidence44
The bandwidth gap is 4.4x and the capacity gap is 4x, which is why these two boxes are not really competing. One decides whether a model fits; the other decides whether it is usable.
Reality
- Evidence38
- Adoption32
- Hype gap+25
- Incentives42
- Confidence35
NVIDIA Dynamo keeps a pre-warmed engine on the same GPUs and hands it the resident weights instead of reloading them. The measured recovery window falls to about 2.6% of a cold restart.
Reality
- Evidence52
- Adoption22
- Hype gap+22
- Incentives86
- Confidence45
SemiAnalysis's AgentX replays recorded coding-agent sessions instead of fixed prompts. On that traffic, the reported Nvidia-AMD cost gap grows as the interactivity target rises.
Reality
- Evidence46
- Adoption24
- Hype gap+34
- Incentives58
- Confidence44
SemiAnalysis has pushed fixed-length serving into maintenance mode. On replayed Claude Code traffic, NVIDIA's GB300 is credited with 15x Hopper, against 40x over H200 on the retired static test.
Reality
- Evidence34
- Adoption18
- Hype gap+42
- Incentives88
- Confidence61
The update adds a path selector and a two-tap convolution rather than layers, recovering most of the accuracy that tripling the drafter bought at 15.2% latency, by the vendor's own numbers.
Reality
- Evidence54
- Adoption66
- Hype gap+16
- Incentives74
- Confidence58
A Kubernetes serving guide makes a point most teams skip: the YAML you reviewed is identical across five engines, and everything it hides is what breaks in production.
Reality
- Evidence30
- Adoption28
- Hype gap+8
- Incentives32
- Confidence42