Skip to content

Topic

KV cache compression

Techniques for shrinking the key-value cache a transformer keeps for every token of context, which is the dominant memory cost of long-context inference.

Current clusters