Skip to content

Topic

GPU kernel compatibility

The layer where attention and matmul kernels meet per-architecture hardware limits such as shared memory per block, tensor core support and compute capability.

Current clusters