Build1 publisher2 min readPublished
Meta's 3.2K-line Triton kernel outruns FlashAttention-4 on jagged Blackwell attention
Meta says its 3.2K-line TLX attention kernel beats FlashAttention-4 on jagged B200 shapes by about 13% forward and 50% backward. Both margins come from GEM's broadcast-query workload in bf16, so another team's gain depends on how closely its shapes match.
The Engineer · Build desk

What happened
- GEM packs variable-length user sequences contiguously with an offsets tensor, because padding them to a fixed length can waste up to half the compute.
- FlashAttention-4's CuteDSL kernels run to about 10K lines, roughly three times the size of Meta's TLX kernel.
- Meta has published the kernel on GitHub in its ads_model_kernel_library repository.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Teams planning a hand-written CuteDSL attention kernel for Blackwell can now benchmark a public Triton-level kernel on their own shapes before committing to a 10K-line codebase.
- constraint The backward margin was measured with a broadcast query, so teams with per-sequence queries cannot count on the contended dQ reduction Meta sped up existing in their workload.
- cost Modeling engineers who take over the kernel still have to reason about explicit barriers and warp specialization, logic TLX moves out of the compiler and into their source.
Attention is the slowest kernel in GEM [17]. Meta started from a Triton JFA kernel that was algorithmically correct but left all on-chip data movement and scheduling to the compiler [5]. It had no control over shared-memory allocation or pipeline depth. It had no explicit barriers or warp specialization [5]. A single warp group issued the loads, the softmax and the matmuls, so the tensor cores stalled whenever that group was not issuing an MMA [6]. TLX lets the author split that work across specialized warps and pipeline it by hand [2]. Meta describes the result as turning a compiler-scheduled, memory-bound kernel into a tightly pipelined, warp-specialized one [7].
The backward figure needs the most care before anyone applies it to another model. In Meta's production case a single dense Q is broadcast across every sequence in the batch, so dQ has to be summed across the entire batch [8]. Meta says this makes the dQ epilogue a heavily contended cross-program reduction, and that several of its optimizations exist to make that reduction fast [9]. The post assumes broadcast_q=True unless it says otherwise [10]. A team with per-sequence queries has no contended dQ sum to speed up. I would rerun the backward pass before quoting the ~50% figure for that workload.
The forward number has its own conditions. Meta benchmarked in bfloat16 on B200 [11]. It limits the win to the jagged shapes that matter for GEM, measured against the May 2026 version of FA4 [3]. For the margin to transfer, a workload needs ragged sequences in GEM's packed layout, the same hardware and precision, and the same FA4 baseline. JFA runs FlashAttention directly on the packed Q, K and V tensors and their offsets, and never materializes a padded token [16].
In Meta's context the tradeoff looks right to me. New modeling ideas such as sliding-window and block-sparse attention land in this kernel, so it gets edited often. Meta describes the hand-written CuteDSL or CUDA alternative as hard to read, extend or fuse [13]. Roughly 10K lines against roughly 3.2K is a ratio of about 3.1 [1]. The code is public, so anyone can check the count [14].
The claim that modeling engineers, not just kernel specialists, can read, extend and fuse the kernel comes from Meta [12]. Line counts measure typing. The hard part of a warp-specialized kernel is knowing which warp waits on which barrier. TLX puts barriers, warp specialization, async TMA and MMA, and explicit SMEM and TMEM allocation into the source as first-class primitives [2]. Plain Triton left those decisions to the compiler [5].
What to watch
- An FA4 release after the May 2026 version that narrows or widens the gap on jagged shapes.
- Results from Meta or outside users running the public kernel with broadcast_q=False or dense fixed-length shapes.
- A new attention variant such as sliding window landed in the TLX kernel by a modeling engineer rather than a kernel specialist.