build1 publisher
Quantizing value residuals and casting softmax to FP8 speeds Nunchux's attention kernel 1.46x on an H200
The September 14th VC-Attention paper reports 1.46x on an H200 and 1.59x on a B200 against BF16 FlashAttention-4, measured at the attention call. How much of that reaches a finished clip depends on how much of your step is attention.
Publishers:runtimewire.com
Reality
- Evidence55
- Adoption18
- Hype gap+18
- Incentives70
- Confidence58