build1 publisher
Paging the KV cache into 16-token blocks caps the waste at 960 KB per sequence
One 8,192-token session on a 27B model holds 512 MB of key-value cache, or 64 KB for every token generated. How many of those sessions fit in free VRAM sets serving concurrency, and paging decides the waste.
Publishers:dev.to
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+24
- Incentives45
- Confidence48