build1 publisher
Meta's MXFP8 FlashAttention-4 kernel, tackling TMEM scale-factor constraints, hits up to 1.6x over BF16
Meta has open-sourced an MXFP8 forward and backward pass for FlashAttention-4 that it already runs in production ads training. The reported 1.6x forward gain over BF16 sits well under the 2-4x the block-scaled MMA instruction advertises.
Publishers:pytorch.org
Reality
- Evidence58
- Adoption47
- Hype gap+12
- Incentives55
- Confidence60