Skip to content

Build1 publisher3 min readPublished

Your 2026 GPU Decision Is Arithmetic: Bytes Per Parameter, Times Parameters, Plus Cache

A dev.to guide argues local LLM capacity planning collapses into one napkin equation. Run it first and the hardware shortlist writes itself, tier names and all.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Your 2026 GPU Decision Is Arithmetic: Bytes Per Parameter, Times Parameters, Plus Cache
Generated illustration

What happened

  • A guide published on dev.to states that every model consumes memory proportional to its parameters, and that the number depends on how the weights are stored (bits per parameter).
  • The guide states that the context window, KV cache and the running application all live in the same memory as the weights, making the 'realistic total' more important than raw weight size.
  • The guide describes the calculation as a straightforward napkin equation and says an LLM either fits in your VRAM or it does not, unlike game performance which depends on drivers, resolution and scene complexity.
  • At 16-bit precision, a 7B model needs about 14GB for weights alone.
  • Each step down the precision ladder roughly halves memory while adding a little quantization error.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A guide published on dev.to reduces local LLM capacity planning to a single line of arithmetic: memory consumed is proportional to parameter count, scaled by how many bits you store each weight in, plus the KV cache and the running application, all of which live in the same memory [1][2]. That matters for a 2026 buy because it converts a shopping decision into a pass/fail fit test, and per the guide a model either fits in your VRAM or it does not [3].

Start with the precision ladder. At 16-bit, a 7B model needs about 14GB for weights alone [4]. Each step down in precision roughly halves memory while adding some quantization error [5], which puts the same 7B model near 3.5GB at 4-bit [1] and inside a mainstream 8GB card [6]. The 4-bit GGUF format is the de facto standard for local inference, and the guide treats Q4 as the community baseline, with anything below it worth doing only when the alternative is not running the model at all [7][8].

The interesting arithmetic starts at 70B. At Q4 that is roughly 40GB of weights [9], which works out to about 0.57 bytes per parameter in practice rather than the theoretical 0.5 [2], plus several more gigabytes for a reasonable context window [10]. Against the current consumer flagship, an RTX 5090 with 32GB of GDDR7 [11], the weights alone are about 8GB short before you budget any context [3]. The guide's conclusion follows: the 5090 handles 7B, 14B and 32B comfortably at Q4, but 70B only fits at Q3 or smaller [12]. A single 24GB card rarely runs a 70B comfortably without aggressive quantization or a short context [13].

Above that line there are three routes, and they price differently in money, complexity and speed [14]. Two RTX 4090s give 48GB combined, enough for a 70B at Q4 [15], or 24GB per card [4], but it requires motherboard support, a high-wattage power supply, and a runtime such as vLLM to shard the model [16]. Unified memory buys capacity instead: Apple's Mac Studio M5 Ultra goes to 192GB shared between CPU and GPU [17], six times the 5090's pool [5], and can run 120B-plus models natively [18]. It pays for that in bandwidth, at 819 GB/s against 1,792 GB/s [19], a gap of roughly 2.2x [6] that directly caps tokens per second [20]. The appliance route is NVIDIA's DGX Spark, 128GB of unified memory at $4,699 [21], about $37 per gigabyte of addressable model space [7].

The load-bearing term is the one most guides skip. The KV cache holds every token in your prompt and every token generated so far [22], and the guide is explicit that this is why a "40GB for a 70B" plan falls apart in practice [23]. It scales with context length, so the honest input to the equation is not the model card, it is your real context window under your real workload.

What to watch: whether 32GB stays the consumer ceiling, since that number alone decides the Q4-versus-Q3 question at 70B [12]. And before signing off on multi-GPU, measure your KV cache at production context length. The difference between 40GB of weights and a working system is entirely in that term.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories