Ollama divides the whole prompt by the time it spent computing one token of it
The daemon's prompt_eval_count includes cache reads and prompt_eval_duration does not, so one qwen2.5:7b prompt reported 13,826 tokens a second and 43 tokens a second half a minute apart. One subtraction fixes it.
Reality
- Evidence72
- Adoption18
- Hype gap−5
- Incentives18
- Confidence66