Skip to content

Build1 publisher3 min readPublished

Local LLM speed follows memory bandwidth only while the whole model fits in VRAM

Llama 8B on an RTX 4060 Ti 16GB fell from 42.5 to 3.8 tokens a second with a fifth of the model in system RAM, according to a dev.to benchmark. For a local coding agent, that makes VRAM for weights plus context the first spec to check, ahead of bandwidth.

The Engineer · Build desk

Illustration accompanying Local LLM speed follows memory bandwidth only while the whole model fits in VRAM

What happened

  • The post caps batch-size-1 decoding at memory bandwidth divided by model size: about 94 tokens a second for a 10 GB model on a 936 GB/s RTX 3090.
  • Runtimes such as llama.cpp, Ollama and vLLM split an oversized model between VRAM and system RAM across the PCIe bus without refusing to load it.
  • The KV cache for Llama 3.1 8B grows from about 1.1 GB at 8k tokens to 4.3 GB at 32k and 17 GB at 131k.
  • Qwen 2.5 32B at 32k context needs roughly 25 GB, so a 16 GB RTX 4070 Ti Super cannot run it properly despite 672 GB/s of bandwidth.
  • The author's value pick is a used RTX 3090 24GB at around $650 to $750, chosen because it pairs 24 GB with a 384-bit bus.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Saving about $70 on the 8 GB RTX 4060 Ti costs half the VRAM, and per the post leaves any model above 8B with real context out of memory.
  • contradiction The post's own measurements sit on both sides of its 20x headline, so a buyer cannot budget a partial spill with one multiplier.
  • constraint The post's recommended 24 GB card cannot hold its own 32B-at-32k example, so a 3090 owner running Qwen 2.5 32B has to give up some context.

Each token requires streaming every weight of the model from VRAM through the compute units, and the cores spend most of that time waiting on the memory bus [1]. For a model that fits, the bandwidth ceiling predicts well. An RTX 4060 Ti at 288 GB/s over a 5 GB model has a ceiling near 58 tokens a second, and the post measured 42.5 [7]. The same logic settles the CPU question. According to the post, a 5.8 GHz Core i9 and a liquid loop do nothing here, and a six-core Ryzen 5 is plenty [18].

The spill figure needs more care. The post calls it a 20x bandwidth drop [6], and that is a ratio of ceilings: 936 GB/s on the 3090 against about 50 GB/s for DDR5 is roughly 19x [3]. Its measured runs land elsewhere. On the 3090, 60 to 75 tokens a second in VRAM against 1.5 to 2 from DDR5 [5] is a 30x to 50x gap [2]. The 4060 Ti partial spill works out to about 11x [1].

I think that partial spill is the most useful number in the post, because bandwidth does not explain it. Suppose 4 GB of the 5 GB model stays on the card at 288 GB/s and the spilled gigabyte moves at PCIe 4.0's 31.5 GB/s, the slowest pipe the post names [6][7][11]. That costs about 14 ms plus 32 ms per token, or roughly 22 tokens a second [4]. The benchmark measured 3.8 [8]. At that rate a token takes about 263 ms, so some 217 ms per token sits outside the bandwidth model [5]. The post does not break that time down. Its conclusion holds all the more. The author's position is that almost fitting counts for nothing: either the model and its context fit in VRAM, or it runs at DDR5 speed [10].

These numbers describe one sequence at a time. The post scopes the bandwidth bound to batch size 1 [1], and its weight sizes assume 4-bit quantization such as GGUF Q4_K_M or AWQ, before any KV cache [11]. A developer running one coding agent on a desktop matches that workload. Because the common runtimes split an oversized model across the PCIe bus [9], the first symptom of a model that does not fit is a slow agent.

The cache estimates are the best-built part of the post. They come from a short Python function fed with the layer and head counts in Llama 3.1 8B's Hugging Face config [13]. The bandwidth figures cite Nvidia's spec pages and the PCIe 4.0 and JEDEC DDR5 standards [19]. Anyone can rerun the function with their own model's config. That matters for agents like Cline, Roo Code and Cursor, which keep the system prompt, tool definitions and file chunks in context while the cache grows linearly [12].

Checked against the post's own card pick, its 32B example comes up short. The 25 GB that Qwen 2.5 32B needs at 32k tokens [14] is a gigabyte more than the 3090's 24 GB [6]. With 19.5 GB of weights on the card, 4.5 GB remains for cache, enough for about 26,800 tokens before runtime overhead [7]. The post's roughly $1,650 tier is built for 32B coding models [17].

What to watch

  • A row-by-row table for the 4060 Ti benchmark, or the same split run on a 3090, would show whether the 91 percent loss scales with the spilled fraction or arrives with the first spilled layer.
  • A measured run of Qwen 2.5 32B on the post's roughly $1,650 tier at the full 32k context would settle whether 24 GB holds the example.
  • Used RTX 3090 prices moving outside the $650 to $750 range would change the post's VRAM-per-dollar ranking.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories