BuildNot yet confirmed elsewhere1 publisher2 min readPublished
Ollama's timings block rates a one-token prefill at 3,875 tokens per second
Ollama's OpenAI-compatible endpoint reported 3,875 prompt tokens per second where the real rate was 125, a dev.to author found. Any benchmark that repeats a prompt and reads that field overstates prefill speed by the share served from cache.
The Engineer · Build desk
Bar chart of two prefill rates from one Ollama response. The timings block's prompt_per_second reads 3875 tokens per second. The dev.to author's honest figure, (prompt_tokens - cached_tokens) / prompt_ms = 1 / 8 ms, is 125 tokens per second.
Prefill rate from one response chunk, prompt_ms 8 In tok/s
| Item | Value | Claim |
|---|---|---|
| Timings block prompt_per_second | 3,875 tok/s | 3 |
| dev.to author's honest figure | 125 tok/s | 7 |
What happened
- The timings block rides on the final usage chunk of the OpenAI-compatible endpoint, and has since Ollama 0.35.1.
- prompt_n counts the whole prompt, cached tokens included, while prompt_ms times only the uncached tokens, so the two operands describe different amounts of work.
- The correct figure can be computed from the same chunk: (prompt_tokens - cached_tokens) divided by prompt_ms gives 125 tokens per second in the post's example.
- The native API picked up the same skew in 0.33.3, when prompt_eval_duration began timing only uncached tokens while prompt_eval_count stayed the total.
- The decode rate, predicted_per_second, is unaffected because generated tokens are never cached.
Why it matters
- exposure A loop that sends one prompt repeatedly and averages prompt_per_second takes one accurate sample first and inflated ones after it, so the mean overstates prefill speed.
- constraint The cached count lives in usage, not in timings, so a harness that logs only the timings object cannot recover the real prefill rate afterward.
- decision In our view a harness should compute (prompt_tokens - cached_tokens) / prompt_ms itself and store the runner field beside the model tag.
The field is a division with mismatched operands. According to the post, prompt_n is the whole prompt with cached tokens included, and prompt_ms is the time spent on the uncached ones only [5]. The block divides one by the other: 31 tokens over 8 ms gives the 3,875 [3]. The token that actually ran, one in 8 ms, comes to 125 per second [7]. The reported figure is 31 times that [19]. It is a fast prefill, mostly because it skipped the prefill.
The inflation factor is total prompt tokens over uncached tokens. The author gives it as count / (count - cached) [6]. With 30 of 31 tokens cached the factor is 31 [2]. At 99 percent cached it would be 100 [21]. A cold request has nothing cached, so the factor is 1 and the first number is honest [20]. The post says a long system prompt that matched can give "anything up to a few hundred" [8].
The 31x transfers to your rig only under conditions. The endpoint has to be the OpenAI-compatible one on 0.35.1 or later [1]. The prompt prefix has to still be cached from an earlier request. The cached share has to be near the post's 30 of 31, because the factor follows that share. A workload whose requests share no prefix sees a factor near 1 [20]. The post's numbers come from one 31-token prompt on the author's own daemon [2].
What is new on the compatible endpoint is who does the division. The server now emits the result itself, labelled as a rate [10]. A client reading prompt_per_second gets a finished number, and the cached count sits in a different object of the same chunk [2].
All of this rests on one post. Its author maintains LLMxRay, whose envelope diff shows the stated rate beside the rate over evaluated tokens and turns the stated cell amber on a repeat run [13]. The author says the check is two curl calls on your own machine: send one prompt twice and read the second answer [14]. The post does not report a response from Ollama.
A second comparability problem sits in the same post. Ollama 0.40 made MLX the default on Apple Silicon for supported architectures, so one model tag can now run on two engines [15]. The daemon names the engine in /api/tags and /api/ps, as mlx, llamacpp, or ggml for models pulled before 0.40 [16]. On the author's Mac mini, nomic-embed-text sat on llamacpp while embeddinggemma-2 ran on mlx [17]. The author says a benchmark result that records only the model name is now ambiguous [18].
What to watch
- Whether Ollama changes how prompt_per_second is computed, or documents it, in a release after 0.40.
- Whether a second person reproduces the 31x on another daemon, which would move it from one author's measurement to a confirmed behavior.
- Whether benchmark write-ups for Apple Silicon start recording the runner field next to the model tag.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+5
- Incentives35
- Confidence55
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Since Ollama 0.35.1 the OpenAI-compatible endpoint attaches a timings block to the final usage chunk.
- [2]
In the author's example, the second answer to a repeated prompt, usage shows prompt_tokens 31 and cached_tokens 30 (in prompt_tokens_details), in the same chunk as the timings block.
- [3]
The timings block in that response shows prompt_n 31, prompt_ms 8 and prompt_per_second 3875.
- [4]
One token was evaluated, in 8 milliseconds, yet the block says 3875 tokens per second.
- [5]
prompt_n is the whole prompt, cached tokens included; prompt_ms is the time spent on the uncached tokens only.
- [6]
Dividing prompt_n by prompt_ms gives a rate inflated by count / (count - cached), which is 31x in the example.
- [7]
The honest figure is in the same chunk: (prompt_tokens - cached_tokens) / prompt_ms = 1 / 8 ms = 125 tok/s.
- [8]
On a long system prompt that matched, the inflation can be 'anything up to a few hundred'.
- [9]
The native API acquired the same formula bug in 0.33.3, when prompt_eval_duration started timing only uncached tokens while prompt_eval_count stayed the total.
- [10]
On the compatible endpoint the server now emits the result of that division itself, labelled as a rate.
- [11]
The decode half, predicted_per_second, is fine because generated tokens are never cached.
- [12]
In the example the decode figures were predicted_n 11, predicted_ms 83 and predicted_per_second 132.5.
- [13]
The author maintains LLMxRay, whose Protocol Observatory envelope diff shows prefill time, the rate as the daemon states it, and the rate over evaluated tokens; the stated cell turns amber on the second run of a prompt.
- [14]
The author says the reproduction is two curl calls on your own machine: send one prompt twice and read the second answer.
- [15]
Ollama 0.40 made MLX the default on Apple Silicon for the architectures it supports, so the same model tag can now run on two different engines.
- [16]
The daemon reports the engine as a runner field in /api/tags details and /api/ps: mlx, llamacpp, or ggml for models pulled before 0.40.
- [17]
On the author's Mac mini, nomic-embed-text sat on llamacpp and embeddinggemma-2 on mlx at the same time.
- [18]
A benchmark result that only records the model name is now ambiguous.
- [19]
The reported prompt rate is 31 times the rate over evaluated tokens.
- [20]
With no cached tokens, or no shared prefix between requests, the factor count / (count - cached) equals 1, so a cold request's prompt rate is not inflated.
- [21]
At 99 percent of the prompt cached, the inflation factor is 100.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toFour numbers Ollama 0.40 gives you now, and which ones to trust
1 article · October 11, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.