build1 publisher
NVIDIA's unoptimized DeepSeek-V3 baseline spends 84% of kernel time moving tokens between GPUs
A JAX and Transformer Engine version of the same GB200 run reports 1,068 TFLOPS per GPU, a 10.4x gain. The 84% share bounds what fixing communication alone could have bought, so the expert matmuls got faster too.
Publishers:developer.nvidia.com
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+35
- Incentives85
- Confidence55