Skip to content

project

SageAttention

Quantized attention kernel family for diffusion and language transformers, aimed at low-bit inference speedups.

Known aliases

  • SageAttention2

Relationships

No evidence-backed relationships are recorded.

Current clusters