Skip to content

Build1 publisher2 min readPublished

NVIDIA's RTX Spark PCs fit a 4-bit 70B model in 128GB with about 90GB to spare

NVIDIA's RTX Spark PCs from Lenovo and Acer ship in October with up to 128GB of shared CPU/GPU memory, enough for a 4-bit 70B model. Long-context agents now fit on owned hardware, though a cost comparison with cloud inference waits on system prices.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying NVIDIA's RTX Spark PCs fit a 4-bit 70B model in 128GB with about 90GB to spare
Photo: acer.com

What happened

  • Reports from IFA 2026 and Reuters describe a 20-core Grace CPU paired with a Blackwell-class RTX GPU, rated at up to 1 Petaflop of 4-bit AI compute.
  • A dev.to analysis of the launch puts typical consumer GPUs at 8 to 16GB of dedicated VRAM, against roughly 35 to 40GB for a capable 70B-class model at 4-bit.
  • NVIDIA's Personal AI Router (PAIR) software can spread inference across several RTX PCs on one local network so models and context can outgrow a single box.
  • Reuters covered the launch as NVIDIA's latest push to move AI inference out of the data center and onto personal computers.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • capability Agents over data that cannot leave the building, such as legal or medical drafting, can run a 70B-class model on owned hardware with no cloud data-residency contract to negotiate.
  • constraint The roughly 90GB of context headroom applies only to the top configuration, because the reports quote memory as up to 128GB.
  • constraint Capacity plans have to assume 4-bit weights, since both the 35 to 40GB fit and the 1 Petaflop rating are 4-bit figures.
  • decision Always-on agents such as inbox triage and log monitoring are the first candidates to move local, because their per-token bill runs around the clock.

The 35GB floor checks out. Seventy billion weights at 4 bits, half a byte each, come to 35GB before the runtime takes anything [2]. That is at least 19GB more than a typical consumer card's VRAM [4]. Put a 35 to 40GB model into a 128GB pool and roughly 88 to 93GB remain [3].

That remainder goes to context. According to the dev.to launch post, agents that need large context windows, such as long-document work, multi-step tool loops and memory-rich assistants, stop being cloud-only on this class of hardware [13]. The design deserves credit. CPU and GPU draw from one pool [4], so a large model and a large context window fit on a desk without a data-center card [14]. That pool is 8 to 16 times the VRAM of a typical consumer card [1].

Fit is the first condition for moving an agent off a hosted API. Speed is the second. The reports summarised in the post list core count, GPU class, memory capacity and FP4 compute, and they do not include tokens per second, memory bandwidth, a system price or power draw [3]. The post opens by calling RTX Spark the most important agent infrastructure story of October [9]. Even so, it calls the 1 Petaflop figure "marketing-adjacent" [8].

The post's cost case is that cloud agents bill per token indefinitely, while a local model costs electricity and one hardware purchase [10]. The shape is right. A break-even needs five inputs: the agent's daily token volume, the cloud price for those tokens, the box's sustained throughput, its power draw and its purchase price. A team already on a hosted API has the first two on its invoices [10].

The software side is ready now. Ollama serves open models on consumer hardware, and n8n has nodes that talk to local model endpoints [12].

PAIR is the piece I would test before relying on it. Splitting one inference job across several PCs puts the office network between parts of the model [6].

What to watch

  • Published prices for the Lenovo and Acer RTX Spark systems, which set the hardware side of any break-even against per-token billing.
  • Independent tokens-per-second measurements of a 4-bit 70B model with a long context on a 128GB RTX Spark machine.
  • NVIDIA documentation on how PAIR splits a model across LAN-connected PCs and what link speed it needs.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories