NVIDIA's RTX Spark PCs from Lenovo and Acer ship in October with up to 128GB of shared CPU/GPU memory, enough for a 4-bit 70B model. Long-context agents now fit on owned hardware, though a cost comparison with cloud inference waits on system prices.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+35
- Incentives
- Insufficient
- Confidence30
Salvatore Sanfilippo's open-source ds4 engine runs a short list of large open-weight models locally by compressing their routed experts to about two bits. Even compressed, the supported builds need high-memory Macs or GPU systems that most people do not already own.
Reality
- Evidence55
- Adoption30
- Hype gap+20
- Incentives
- Insufficient
- Confidence60
VIDRAFT says POCKET-Darwin-180B, a 111 GB 4-bit GGUF build of its 180B mixture-of-experts model, runs in llama.cpp on about $1,400 of consumer hardware. The accuracy evidence so far is one MMLU-Pro comparison that VIDRAFT reports itself.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+45
- Incentives60
- Confidence35
Helion's autotuned GEMM beat vLLM's default CUTLASS and DeepGEMM backends on Hopper GPUs, by more than 10% throughput on some workloads, its authors report. The gain rests on per-shape tuning that can run for hours, a cost teams pay in place of kernel maintenance.
Reality
- Evidence35
- Adoption15
- Hype gap+10
- Incentives70
- Confidence40
Lokutor released Oído, open-source transcription of any English sentence on a $5 ESP32-S3, with every figure so far computed off the board. Whether hardware built for command lists can take open-ended speech now depends on how the real boards measure.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+30
- Incentives75
- Confidence40
Gemma 4 decodes on SageMaker's smallest GPU, a T4, at 0.8x an L4's speed with matching outputs from a patched vLLM image, a dev.to benchmark reports. The T4 is the cheaper choice per token only when it rents for under 80% of the L4's hourly rate.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives35
- Confidence40
Llama 8B on an RTX 4060 Ti 16GB fell from 42.5 to 3.8 tokens a second with a fifth of the model in system RAM, according to a dev.to benchmark. For a local coding agent, that makes VRAM for weights plus context the first spec to check, ahead of bandwidth.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence40
Qwen Flash-Next NVFP4 ran in vLLM at 131,072 tokens of context and 16 sequences only after loader patches and cuts to its original targets. Its one load test covers only the bfloat16 KV-cache baseline, so operators on the later B12x stack have to run their own.
Reality
- Evidence45
- Adoption8
- Hype gap0
- Incentives
- Insufficient
- Confidence40
Developer xbill9 rebuilt Google's QAT Gemma 4 26B as int4 and fit it on one TPU v6e with 53,888 tokens of KV cache, against 3,456 for RedHat's FP8 build. Throughput nearly doubles, with accuracy checked on one classification suite.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+15
- Incentives40
- Confidence55
The weights formula is the easy part of sizing local inference. Quantization metadata, the KV cache and the runtime's own buffers decide whether a 70B model fits, and a dev.to walkthrough shows where the advertised bit width stops helping.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives20
- Confidence55
Ternary weights at 1.76 bits put a 27B model into a 5.9GB file and let a laptop decode it at 28.1 tokens a second. The retention figure comes from Prism ML's own benchmark suite, not the table on the model card.
Reality
- Evidence38
- Adoption42
- Hype gap+34
- Incentives74
- Confidence46
Rohan Pinto says BitNet b1.58 runs a 100-billion-parameter model at five to seven tokens a second on a single CPU. That is enough for overnight batches and air-gapped servers, and short of what a cloud tier needs.
Reality
- Evidence22
- Adoption
- Insufficient
- Hype gap+58
- Incentives66
- Confidence57
AWS counted 13 SageMaker inference launches so far in 2026. The one that changes production behaviour most is a prioritized list of up to five instance types, each allowed its own model optimization settings.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives85
- Confidence55
PrismML's ternary build of Qwen3.8 27B keeps 98.2 percent of the full-precision benchmark average on both of the company's inconsistent scorecards, and the loss it does take is concentrated in knowledge and reasoning.
Reality
- Evidence38
- Adoption20
- Hype gap+32
- Incentives72
- Confidence55
The ONNX export alone ran 1.58x faster than eager PyTorch with bit-identical accuracy. The int8 pass after it cost 8% in latency and bought a 64 MB download, small enough to serve as a static page with no backend.
Reality
- Evidence58
- Adoption10
- Hype gap−10
- Incentives30
- Confidence55
A dev.to guide to running local models on 8GB prices the KV cache between 15KB and 160KB per token depending on architecture. At 32K tokens held, that spread is the difference between 0.5GB and 5GB of a fixed budget.
Reality
- Evidence22
- Adoption
- Insufficient
- Hype gap+45
- Incentives62
- Confidence58
A dev.to walkthrough argues that the PROCESSOR field and the startup inference compute log line narrow a slow Ollama to about five named causes, and each of the three states PROCESSOR can print points at a different fix.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives30
- Confidence55
Two arms on the same laptop differ by one flag. The small card wins because llama.cpp leaves Gemma 4's 1.93 GB per-layer embedding table in mmap and pulls a few rows per token, so only about 1.08 GB of body is resident.
Reality
- Evidence64
- Adoption12
- Hype gap+5
- Incentives20
- Confidence58
One endpoint fronts more than 300 models, but the company running the GPUs picks the inference engine and the quantization. A dev.to writeup says the quality gap that follows turns up in the response body, while the status code still reads success.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+25
- Incentives
- Insufficient
- Confidence45
Alibaba's 27B scores 52 on the Artificial Analysis index from a 17GB quantized file. Filling its 262,144-token window needs roughly 16 GiB of KV cache on top of that, so the file size is the smaller half of the sizing question.
Reality
- Evidence30
- Adoption15
- Hype gap+45
- Incentives50
- Confidence32
Earlier coverage
- Chip Huyen puts inference at 10 to 100 times a model's training compute
Build · September 13, 2026 · 1 publisher
- An unquantized model on the main thread cost a health app its screening feature
Build · September 13, 2026 · 1 publisher
- Edge0's 35B would need 22 GB/s of SSD bandwidth without the file cache
Build · September 11, 2026 · 2 publishers
- Bartowski broke tensors one at a time to find where GGUF bits belong
Build · September 11, 2026 · 1 publisher
- 16 KB of per-row scales recover 2.2 of the 2.5 points int4 costs an embedding table
Build · September 10, 2026 · 1 publisher
- NVFP4 squeezes Qwen3.8's 2.4 trillion weights onto eight B300s at 150 GB a GPU
Build · September 9, 2026 · 1 publisher
- llama.cpp takes roughly half an hour to reach first token on an RTX 5090
Build · September 8, 2026 · 1 publisher
- Size the model to the RAM you own before the 45-minute download
Build · September 6, 2026 · 1 publisher
- One registry entry gates 16 of FreeToolHub's 21 in-browser tools on a single model
Build · September 5, 2026 · 1 publisher
- Before you buy another GPU, check num_ctx and the rope base
Build · August 22, 2026 · 1 publisher
- Unsloth's 10% quant claim is really about which machines can run a 27B model
Build · August 19, 2026 · 1 publisher
- NVIDIA's 4-bit Nemotron shows what aggressive quantization costs: you have to retrain for it
Build · August 17, 2026 · 1 publisher
- Your 2026 GPU Decision Is Arithmetic: Bytes Per Parameter, Times Parameters, Plus Cache
Build · August 16, 2026 · 1 publisher
- A 27B Apache-2.0 model in 17GB makes local inference a wiring decision, not a demo
Build · August 15, 2026 · 1 publisher