Skip to content

Topic

Attention kernels

Fused implementations of transformer attention that tile the computation and keep intermediates in on-chip memory instead of materialising the full score matrix.

Current clusters