Ollama's timings block rates a one-token prefill at 3,875 tokens per second
Ollama's OpenAI-compatible endpoint reported 3,875 prompt tokens per second where the real rate was 125, a dev.to author found. Any benchmark that repeats a prompt and reads that field overstates prefill speed by the share served from cache.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+5
- Incentives35
- Confidence55