build1 distinct publisher
Prefill ate 85% of a 291-second answer, and the fix was a dedup key and a cache slot
One question over nine files assembled a 13,773-token prompt, and llama.cpp reused three tokens of it. Reading beat writing by about five to one.
Publishers:dev.to
Reality
- Evidence58
- Adoption22
- Hype gap+12
- Incentives32