Skip to content

other

Multi-Head Latent Attention

Attention design that caches one low-dimensional latent vector per token and reconstructs per-head keys and values from it, keeping positional information in a separate rotary stream.

Current clusters