build1 distinct publisher
Cached prefixes push 91% of a ReAct agent's LLM time into decode
An arXiv tracing study of Claude Code agents on Gemma and Qwen measured prefix-cache hit rates between 84.6 and 99.5 percent, which moves the serving bottleneck to how long you can keep KV blocks resident between tool calls.
Publishers:arxiv.org
Reality
- Evidence58
- Adoption25
- Hype gap+12
- Incentives
- Insufficient
- Confidence