Apple has not confirmed the plan and the M8 Ultra is years from release. Seven other companies have already licensed the same rack interconnect from Nvidia, and their dates are the ones a buyer can use.
Reality
- Evidence45
- Adoption50
- Hype gap+20
- Incentives70
- Confidence55
MLPerf Inference v6.1 preview submissions put Vera Rubin NVL72 at up to 3.7x GB300 on Qwen3-VL and up to 2.5x on DeepSeek-R1, on two different inference frameworks. The four-rack 99% scaling result is an offline number.
Reality
- Evidence45
- Adoption30
- Hype gap+35
- Incentives85
- Confidence55
Qwen3.8-Flash-Next puts 36 Gated DeltaNet layers and 12 sparse-attention layers on Hugging Face, which means the retrieval budget Qwen4 will inherit is something you can measure against your own traces now.
Perspective Coverage
5 publishers
- Builder
- Builder 52%
- Operator
- Operator 28%
- Investor
- Investor 20%
Reality
- Evidence58
- Adoption52
- Hype gap+32
- Incentives76
- Confidence71
SemiAnalysis has pushed fixed-length serving into maintenance mode. On replayed Claude Code traffic, NVIDIA's GB300 is credited with 15x Hopper, against 40x over H200 on the retired static test.
Reality
- Evidence34
- Adoption18
- Hype gap+42
- Incentives88
- Confidence61
NVIDIA says Alibaba's largest open-weight model serves over 4K tokens/sec/GPU and 350 tokens/sec/user in FP8 on a GB300 NVL72. That figure is the self-hosting floor, not a benchmark.
Reality
- Evidence32
- Adoption42
- Hype gap+34
- Incentives88
- Confidence44