build1 publisher
Kimi Linear's 75% KV cache cut comes from making three of every four layers recurrent
Moonshot AI's Kimi Linear reports 75% less KV cache and 6.3x faster decoding than MLA at 1M tokens, using three linear layers per attention layer. The speedup falls to parity at 4k tokens, so the saving goes to traffic that runs at hundreds of thousands of tokens.
Publishers:dev.to
Reality
- Evidence40
- Adoption35
- Hype gap+25
- Incentives60
- Confidence45