Skip to content

BuildNot yet confirmed elsewhere1 publisher2 min readPublished

Ollama's timings block rates a one-token prefill at 3,875 tokens per second

Ollama's OpenAI-compatible endpoint reported 3,875 prompt tokens per second where the real rate was 125, a dev.to author found. Any benchmark that repeats a prompt and reads that field overstates prefill speed by the share served from cache.

The Engineer · Build desk

How we use AISend a correction

Ollama field says 3,875 tok/s; dev.to author computes 125 Prefill rate in one Ollama response: the timings block's prompt_per_second versus the dev.to author's (prompt_tokens - cached_tokens) / prompt_ms.

Bar chart of two prefill rates from one Ollama response. The timings block's prompt_per_second reads 3875 tokens per second. The dev.to author's honest figure, (prompt_tokens - cached_tokens) / prompt_ms = 1 / 8 ms, is 125 tokens per second.

Ollama field says 3,875 tok/s; dev.to author computes 125 (Prefill rate from one response chunk, prompt_ms 8)
ItemValueClaim
Timings block prompt_per_second3,875 tok/s3
dev.to author's honest figure125 tok/s7

What happened

  • The timings block rides on the final usage chunk of the OpenAI-compatible endpoint, and has since Ollama 0.35.1.
  • prompt_n counts the whole prompt, cached tokens included, while prompt_ms times only the uncached tokens, so the two operands describe different amounts of work.
  • The correct figure can be computed from the same chunk: (prompt_tokens - cached_tokens) divided by prompt_ms gives 125 tokens per second in the post's example.
  • The native API picked up the same skew in 0.33.3, when prompt_eval_duration began timing only uncached tokens while prompt_eval_count stayed the total.
  • The decode rate, predicted_per_second, is unaffected because generated tokens are never cached.

Why it matters

  • exposure A loop that sends one prompt repeatedly and averages prompt_per_second takes one accurate sample first and inflated ones after it, so the mean overstates prefill speed.
  • constraint The cached count lives in usage, not in timings, so a harness that logs only the timings object cannot recover the real prefill rate afterward.
  • decision In our view a harness should compute (prompt_tokens - cached_tokens) / prompt_ms itself and store the runner field beside the model tag.

The field is a division with mismatched operands. According to the post, prompt_n is the whole prompt with cached tokens included, and prompt_ms is the time spent on the uncached ones only [5]. The block divides one by the other: 31 tokens over 8 ms gives the 3,875 [3]. The token that actually ran, one in 8 ms, comes to 125 per second [7]. The reported figure is 31 times that [19]. It is a fast prefill, mostly because it skipped the prefill.

The inflation factor is total prompt tokens over uncached tokens. The author gives it as count / (count - cached) [6]. With 30 of 31 tokens cached the factor is 31 [2]. At 99 percent cached it would be 100 [21]. A cold request has nothing cached, so the factor is 1 and the first number is honest [20]. The post says a long system prompt that matched can give "anything up to a few hundred" [8].

The 31x transfers to your rig only under conditions. The endpoint has to be the OpenAI-compatible one on 0.35.1 or later [1]. The prompt prefix has to still be cached from an earlier request. The cached share has to be near the post's 30 of 31, because the factor follows that share. A workload whose requests share no prefix sees a factor near 1 [20]. The post's numbers come from one 31-token prompt on the author's own daemon [2].

What is new on the compatible endpoint is who does the division. The server now emits the result itself, labelled as a rate [10]. A client reading prompt_per_second gets a finished number, and the cached count sits in a different object of the same chunk [2].

All of this rests on one post. Its author maintains LLMxRay, whose envelope diff shows the stated rate beside the rate over evaluated tokens and turns the stated cell amber on a repeat run [13]. The author says the check is two curl calls on your own machine: send one prompt twice and read the second answer [14]. The post does not report a response from Ollama.

A second comparability problem sits in the same post. Ollama 0.40 made MLX the default on Apple Silicon for supported architectures, so one model tag can now run on two engines [15]. The daemon names the engine in /api/tags and /api/ps, as mlx, llamacpp, or ggml for models pulled before 0.40 [16]. On the author's Mac mini, nomic-embed-text sat on llamacpp while embeddinggemma-2 ran on mlx [17]. The author says a benchmark result that records only the model name is now ambiguous [18].

What to watch

  • Whether Ollama changes how prompt_per_second is computed, or documents it, in a release after 0.40.
  • Whether a second person reproduces the 31x on another daemon, which would move it from one author's measurement to a confirmed behavior.
  • Whether benchmark write-ups for Apple Silicon start recording the runner field next to the model tag.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence55
Adoption
Insufficient
Hype gap+5
Incentives35
Confidence55
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Since Ollama 0.35.1 the OpenAI-compatible endpoint attaches a timings block to the final usage chunk.

    ReportedSupportedSource: dev.to post by the maintainer of LLMxRayView cited source
  2. [2]

    In the author's example, the second answer to a repeated prompt, usage shows prompt_tokens 31 and cached_tokens 30 (in prompt_tokens_details), in the same chunk as the timings block.

    ReportedSupportedSource: dev.to post by the maintainer of LLMxRayView cited source
  3. [3]

    The timings block in that response shows prompt_n 31, prompt_ms 8 and prompt_per_second 3875.

    ReportedSupportedSource: dev.to post by the maintainer of LLMxRayView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. dev.to

    1 article · October 11, 2026

    Four numbers Ollama 0.40 gives you now, and which ones to trust

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories