Skip to content

Topic

GPU kernel authoring

The practice of writing the low-level compute routines that machine-learning frameworks call for matmuls, attention and reductions, in languages and DSLs such as CUDA, Triton and Helion.

Current clusters