NVIDIA measured TensorRT LLM holding 96.1 to 98.2 percent of its non-confidential output throughput on Blackwell, and it got there by unpinning host memory on the affected paths, moving decode readback off the scheduler thread, and timing kernel tactics with the GPU's global timer instead of CUDA events.
Reality
- Evidence58
- Adoption22
- Hype gap+10
- Incentives80
- Confidence55
DataEnclave uses Nvidia confidential computing so a model developer's base weights and a customer's fine-tuned weights can run in one enclave under separate key control. Vast says it charges nothing extra for it.
Reality
- Evidence38
- Adoption15
- Hype gap+30
- Incentives80
- Confidence45
MLPerf Inference v6.1 preview submissions put Vera Rubin NVL72 at up to 3.7x GB300 on Qwen3-VL and up to 2.5x on DeepSeek-R1, on two different inference frameworks. The four-rack 99% scaling result is an offline number.
Reality
- Evidence45
- Adoption30
- Hype gap+35
- Incentives85
- Confidence55
The Intel spinout's Active Compute Fabric targets Nvidia's NVLink but reaches customers only through platform builders such as Qualcomm, on terms neither company has described.
Reality
- Evidence55
- Adoption32
- Hype gap+35
- Incentives72
- Confidence60
KeyBanc's optical thesis for Marvell and Marvell's own revenue guidance both land near $30 billion, four years apart and measuring different things. The purchase price for the technology is already committed.
Reality
- Evidence32
- Adoption45
- Hype gap+30
- Incentives62
- Confidence36
The Windows on Arm port and the compute-capability-107 Rubin preview are the headline items, but the two lines that touch a running cluster are the unbundled driver installer and the new CDMM default on coherent platforms.
Reality
- Evidence60
- Adoption20
- Hype gap+18
- Incentives82
- Confidence58
The quarter came in at $96.2B. The argument that compute must now leave the mega-cluster rests on one analyst's own 30 GW estimate, and on nothing Nvidia itself said.
Reality
- Evidence32
- Adoption20
- Hype gap+46
- Incentives62
- Confidence54
NVIDIA says Alibaba's largest open-weight model serves over 4K tokens/sec/GPU and 350 tokens/sec/user in FP8 on a GB300 NVL72. That figure is the self-hosting floor, not a benchmark.
Reality
- Evidence32
- Adoption42
- Hype gap+34
- Incentives88
- Confidence44