Salvatore Sanfilippo's open-source ds4 engine runs a short list of large open-weight models locally by compressing their routed experts to about two bits. Even compressed, the supported builds need high-memory Macs or GPU systems that most people do not already own.
Reality
- Evidence55
- Adoption30
- Hype gap+20
- Incentives
- Insufficient
- Confidence60
VIDRAFT says POCKET-Darwin-180B, a 111 GB 4-bit GGUF build of its 180B mixture-of-experts model, runs in llama.cpp on about $1,400 of consumer hardware. The accuracy evidence so far is one MMLU-Pro comparison that VIDRAFT reports itself.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+45
- Incentives60
- Confidence35
Prism's ternary Bonsai 2 27B ran at two to three tokens a second on a CPU-only Hetzner VPS in a dev.to test, against Simon Willison's 20 to 44 on a Mac. The sub-6GB file fits a 16GB box easily, but at that speed it only suits batch jobs.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+35
- Incentives
- Insufficient
- Confidence45
Llama 8B on an RTX 4060 Ti 16GB fell from 42.5 to 3.8 tokens a second with a fifth of the model in system RAM, according to a dev.to benchmark. For a local coding agent, that makes VRAM for weights plus context the first spec to check, ahead of bandwidth.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence40
Developer xbill9 rebuilt Google's QAT Gemma 4 26B as int4 and fit it on one TPU v6e with 53,888 tokens of KV cache, against 3,456 for RedHat's FP8 build. Throughput nearly doubles, with accuracy checked on one classification suite.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+15
- Incentives40
- Confidence55
LM Studio 0.4.0 shipped llmster, a headless daemon that runs it on the Linux GPU servers where MIT-licensed Ollama already worked. Teams whose policy demands auditable source now decide on licence, since only LM Studio's lms CLI carries an MIT grant.
Reality
- Evidence45
- Adoption35
- Hype gap+10
- Incentives20
- Confidence45
Ternary weights at 1.76 bits put a 27B model into a 5.9GB file and let a laptop decode it at 28.1 tokens a second. The retention figure comes from Prism ML's own benchmark suite, not the table on the model card.
Reality
- Evidence38
- Adoption42
- Hype gap+34
- Incentives74
- Confidence46
The demo uploads a model file as base64 chunks over GET, then starts an inference server on it. It shows why egress and WAF rules keyed on the HTTP verb miss what the URL is doing.
Reality
- Evidence64
- Adoption9
- Hype gap+14
- Incentives30
- Confidence56
PrismML's ternary build of Qwen3.8 27B keeps 98.2 percent of the full-precision benchmark average on both of the company's inconsistent scorecards, and the loss it does take is concentrated in knowledge and reasoning.
Reality
- Evidence38
- Adoption20
- Hype gap+32
- Incentives72
- Confidence55
Every local runner now reads the same GGUF file, so the binding decision is the quant tag and the gigabyte or two of context that has to fit beside it. Ollama sets the GPU offload itself; llama.cpp lets you set it.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+24
- Incentives44
- Confidence45
A dev.to walkthrough argues that the PROCESSOR field and the startup inference compute log line narrow a slow Ollama to about five named causes, and each of the three states PROCESSOR can print points at a different fix.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives30
- Confidence55
Two arms on the same laptop differ by one flag. The small card wins because llama.cpp leaves Gemma 4's 1.93 GB per-layer embedding table in mmap and pulls a few rows per token, so only about 1.08 GB of body is resident.
Reality
- Evidence64
- Adoption12
- Hype gap+5
- Incentives20
- Confidence58
Alibaba's 27B scores 52 on the Artificial Analysis index from a 17GB quantized file. Filling its 262,144-token window needs roughly 16 GiB of KV cache on top of that, so the file size is the smaller half of the sizing question.
Reality
- Evidence30
- Adoption15
- Hype gap+45
- Incentives50
- Confidence32
A dev.to guide puts the llama.cpp and Ollama decision on ownership. Under Ollama the named model is the unit you operate; under llama.cpp it is the llama-server process and the seven flags in its command line.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap0
- Incentives30
- Confidence58
ROCm sits under PyTorch, vLLM and SGLang; Vulkan is what llama.cpp-class engines compile shaders against. The engine you already run narrows the choice to one, and the rest of the work is reading logs.
Reality
- Evidence52
- Adoption22
- Hype gap−8
- Incentives30
- Confidence48
Per-tensor layout maps now drive his GGUF releases. The sensitivity data behind them came out of more than 1,000 quantizations of two Qwen3.5 models, scored by KL divergence against bf16 on wikitext-2-raw.
Reality
- Evidence57
- Adoption28
- Hype gap+8
- Incentives42
- Confidence55
A 15M-parameter model streams English text on a 2007 PSP at about one token per second. That is the extreme end of a sizing rule. The harder half of that rule is checking whether the file that fits is a format its own maintainer recommends.
Reality
- Evidence38
- Adoption31
- Hype gap+12
- Incentives58
- Confidence46
Treating the runtime and the weights as a container image buys you a hardened default docker run and a signable artifact. On Apple Silicon, though, that same boundary costs you the GPU. One hands-on run puts a number on that cost.
Reality
- Evidence52
- Adoption20
- Hype gap−5
- Incentives30
- Confidence48
Single-user inference streams weights out of memory, so the sizing question for a private document assistant is RAM and prompt length rather than which accelerator a vendor quoted. The laptop in question was three years old.
Reality
- Evidence46
- Adoption14
- Hype gap+9
- Incentives22
- Confidence41
A capacity model for streaming a mixture-of-experts checkpoint agreed with a stranger's conversion to within 1.1%, but that agreement was two errors of opposite sign, and the resolved answer was sitting in GGUF metadata on the author's own disk.
Reality
- Evidence46
- Adoption
- Insufficient
- Hype gap−10
- Incentives22
- Confidence57
Earlier coverage
- Binding a local model server to 0.0.0.0 hands the LAN an unauthenticated API
Build · August 30, 2026 · 1 publisher
- Ollama, vLLM, SGLang: the throughput ceiling is set by the queue, not the weights
Build · August 22, 2026 · 1 publisher
- "Local" Is A Statement About Inference, Not About Sockets
Build · August 20, 2026 · 1 publisher
- Unsloth's 10% quant claim is really about which machines can run a 27B model
Build · August 19, 2026 · 1 publisher
- Ornith-1.0's benchmarks are fine. Ollama can't parse its tool calls.
Build · August 18, 2026 · 1 publisher
- You Procured Qwen. Your Edge Boxes Are Running Somebody Else's File.
Leadership · August 18, 2026 · 1 publisher
- A refusal-stripped 27B model now ships as a 17.9 GB llama.cpp pull
Build · August 16, 2026 · 1 publisher
- Your 2026 GPU Decision Is Arithmetic: Bytes Per Parameter, Times Parameters, Plus Cache
Build · August 16, 2026 · 1 publisher
- A .keras config can carry a marshalled Python code object, and load_model runs it
Build · August 16, 2026 · 1 publisher
- A 27B Apache-2.0 model in 17GB makes local inference a wiring decision, not a demo
Build · August 15, 2026 · 1 publisher