Salvatore Sanfilippo's open-source ds4 engine runs a short list of large open-weight models locally by compressing their routed experts to about two bits. Even compressed, the supported builds need high-memory Macs or GPU systems that most people do not already own.
Reality
- Evidence55
- Adoption30
- Hype gap+20
- Incentives
- Insufficient
- Confidence60
Cloudflare released two Jev-API-compatible decision models, Clef and Clef-flash, on Workers AI and as Apache 2.0 weights on Hugging Face. Typed classification steps in agent code can now move between providers or onto owned hardware, as long as they stay inside the text-only, 32k-context features Jev supports.
Perspective Coverage
3 publishers
- Builder
- Builder 53%
- Operator
- Operator 27%
- Investor
- Investor 20%
Reality
- Evidence50
- Adoption18
- Hype gap+35
- Incentives70
- Confidence60
llama-server left the request hanging or crashed outright in all 73 trials where a 17,653-token prompt arrived 3 to 46 ms before its sleep timer fired. A one-token warm-up request sent first avoided both failures in 123 of 123 tries.
Reality
- Evidence72
- Adoption
- Insufficient
- Hype gap+5
- Incentives25
- Confidence66
VIDRAFT says POCKET-Darwin-180B, a 111 GB 4-bit GGUF build of its 180B mixture-of-experts model, runs in llama.cpp on about $1,400 of consumer hardware. The accuracy evidence so far is one MMLU-Pro comparison that VIDRAFT reports itself.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+45
- Incentives60
- Confidence35
llama.cpp merged a /v1/systemone endpoint on October 2 that returns typed answers with probabilities from five open model families running locally. Whether it can replace hosted classification calls depends on calibration that each team has to measure on its own data.
Perspective Coverage
3 publishers
- Builder
- Builder 48%
- Operator
- Operator 29%
- Investor
- Investor 23%
Reality
- Evidence50
- Adoption15
- Hype gap+5
- Incentives55
- Confidence55
NobodyWho rebuilt the core of TypeSafe AI's Jev with a local 0.6B Qwen model and 25 lines of Python. Teams pricing Jev for routing or tool-call safety checks can test that local baseline first, provided they measure its calibration on their own labeled decisions.
Reality
- Evidence55
- Adoption55
- Hype gap+40
- Incentives60
- Confidence55
Prism's ternary Bonsai 2 27B ran at two to three tokens a second on a CPU-only Hetzner VPS in a dev.to test, against Simon Willison's 20 to 44 on a Mac. The sub-6GB file fits a 16GB box easily, but at that speed it only suits batch jobs.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+35
- Incentives
- Insufficient
- Confidence45
Llama 8B on an RTX 4060 Ti 16GB fell from 42.5 to 3.8 tokens a second with a fifth of the model in system RAM, according to a dev.to benchmark. For a local coding agent, that makes VRAM for weights plus context the first spec to check, ahead of bandwidth.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence40
Developer xbill9 rebuilt Google's QAT Gemma 4 26B as int4 and fit it on one TPU v6e with 53,888 tokens of KV cache, against 3,456 for RedHat's FP8 build. Throughput nearly doubles, with accuracy checked on one classification suite.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+15
- Incentives40
- Confidence55
LM Studio 0.4.0 shipped llmster, a headless daemon that runs it on the Linux GPU servers where MIT-licensed Ollama already worked. Teams whose policy demands auditable source now decide on licence, since only LM Studio's lms CLI carries an MIT grant.
Reality
- Evidence45
- Adoption35
- Hype gap+10
- Incentives20
- Confidence45
Liquid AI shipped a 279.5-million-parameter draft model for LFM2.5-VL-3B on September 24, reporting 2.30x to 3.13x faster decoding on an Apple M5 Max. Image encoding and prompt prefill are unchanged, and the company's tests did not cover quantized deployments.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+15
- Incentives70
- Confidence60
Ternary weights at 1.76 bits put a 27B model into a 5.9GB file and let a laptop decode it at 28.1 tokens a second. The retention figure comes from Prism ML's own benchmark suite, not the table on the model card.
Reality
- Evidence38
- Adoption42
- Hype gap+34
- Incentives74
- Confidence46
The engine's rendered frame is the fixed input, and a studio decides what the model may touch using per-pixel masks and two intensity parameters. NBA 2K27 ships it first, on RTX 50 Series GPUs only.
Reality
- Evidence42
- Adoption24
- Hype gap+18
- Incentives88
- Confidence50
GLM-5.3-Flash leads agentic terminal work, DeepSeek V4 Flash is billed as the cheapest per token, and a 2.52B MiniCPM5-2B runs locally under Apache 2.0. The comparison flags most of those numbers as vendor-reported.
Reality
- Evidence34
- Adoption27
- Hype gap+26
- Incentives58
- Confidence41
The demo uploads a model file as base64 chunks over GET, then starts an inference server on it. It shows why egress and WAF rules keyed on the HTTP verb miss what the URL is doing.
Reality
- Evidence64
- Adoption9
- Hype gap+14
- Incentives30
- Confidence56
PrismML's ternary build of Qwen3.8 27B keeps 98.2 percent of the full-precision benchmark average on both of the company's inconsistent scorecards, and the loss it does take is concentrated in knowledge and reasoning.
Reality
- Evidence38
- Adoption20
- Hype gap+32
- Incentives72
- Confidence55
Every local runner now reads the same GGUF file, so the binding decision is the quant tag and the gigabyte or two of context that has to fit beside it. Ollama sets the GPU offload itself; llama.cpp lets you set it.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+24
- Incentives44
- Confidence45
A dev.to guide to running local models on 8GB prices the KV cache between 15KB and 160KB per token depending on architecture. At 32K tokens held, that spread is the difference between 0.5GB and 5GB of a fixed budget.
Reality
- Evidence22
- Adoption
- Insufficient
- Hype gap+45
- Incentives62
- Confidence58
An Apache-2.0 memory layer for coding agents returns a candidate only when several signals agree, and it labels every answer STRONG, WEAK or MISS so the calling agent has to branch on confidence before reusing an old fix.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+15
- Incentives65
- Confidence52
A dev.to walkthrough argues that the PROCESSOR field and the startup inference compute log line narrow a slow Ollama to about five named causes, and each of the three states PROCESSOR can print points at a different fix.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives30
- Confidence55
Earlier coverage
- NVFP4 and cache reuse cut MLPerf's edge agent workload to 24 minutes on one Jetson Thor
Build · September 16, 2026 · 1 publisher
- A 4 GB laptop GPU decodes quantised Gemma 4 at 4.27x the CPU rate on 1598 MiB
Build · September 16, 2026 · 1 publisher
- Filling Qwen 3.8 27B's native context costs about as much memory as its weights
Build · September 15, 2026 · 1 publisher
- Running llama-server puts context, KV cache and GPU placement in your command line
Build · September 14, 2026 · 1 publisher
- A diffusion drafter lost to Gemma's own Assistant model on a 12GB RTX 3060
Build · September 14, 2026 · 1 publisher
- A million serverless briefings for $48 implies a Lambda rate of $0.000004 per GB-second
Build · September 12, 2026 · 1 publisher
- A backend that detects the AMD GPU can still leave operations on the CPU
Build · September 12, 2026 · 1 publisher
- 6,935 exposed Ollama servers answered an internet scan without asking for credentials
Security · September 11, 2026 · 1 publisher
- Bartowski broke tensors one at a time to find where GGUF bits belong
Build · September 11, 2026 · 1 publisher
- Overnight laptop runs took over most of one Rust developer's Opus coding work
Build · September 10, 2026 · 1 publisher
- A 6% driver reserve decides which models fit on a $2,000 pair of P40s
Build · September 8, 2026 · 1 publisher
- llama.cpp takes roughly half an hour to reach first token on an RTX 5090
Build · September 8, 2026 · 1 publisher
- Size the model to the RAM you own before the 45-minute download
Build · September 6, 2026 · 1 publisher
- NVIDIA's edge-agent case rests on compact open models matching data-center capabilities
Build · September 4, 2026 · 1 publisher
- RamaLama ships models as OCI images you can inspect and sign
Build · August 31, 2026 · 1 publisher
- A 7B model at 11 tokens per second cleared the bar for private contract Q&A
Build · August 30, 2026 · 1 publisher
- Reading one ambiguous config key correctly pushed a passing cost model to 3.4% error
Build · August 30, 2026 · 1 publisher
- Running the model on the laptop turns a subscription line into a maintenance chore
Product · August 29, 2026 · 1 publisher
- Four-bit weights leave 6 GB on a 24 GB card for KV cache and vision tensors
Build · August 29, 2026 · 1 publisher
- Intel puts its Arc GPU operating knowledge inside the coding agent already installed
Build · August 28, 2026 · 1 publisher
- Prefill ate 85% of a 291-second answer, and the fix was a dedup key and a cache slot
Build · August 26, 2026 · 1 publisher
- Mac Studio M5 Ultra vs DGX Spark: capacity says what fits, bandwidth says what you wait for
Build · August 25, 2026 · 1 publisher
- Parallels ships OpenGL 4.3 on Apple silicon, and the fleet question moves to which chip you own
Product · August 25, 2026 · 1 publisher
- Pi ships four tools and no sandbox, so the guardrails are on your build list
Build · August 23, 2026 · 1 publisher
- A 284B model at 25 tokens a second on one 5090, and 192 GiB of DDR5 doing the quiet part
Build · August 22, 2026 · 1 publisher
- Base Compute hands kernel tuning to agents; the carryover claim is the unmeasured part
Build · August 21, 2026 · 1 publisher
- The flash_attn error in llama.cpp is a layout constraint, and it decides your context window
Build · August 21, 2026 · 1 publisher
- Unsloth's 10% quant claim is really about which machines can run a 27B model
Build · August 19, 2026 · 1 publisher
- Inco AI's DFlash 2: 21% longer accepted drafts for 1.3% latency and 18.5M parameters
Build · August 19, 2026 · 1 publisher
- You Procured Qwen. Your Edge Boxes Are Running Somebody Else's File.
Leadership · August 18, 2026 · 1 publisher
- A refusal-stripped 27B model now ships as a 17.9 GB llama.cpp pull
Build · August 16, 2026 · 1 publisher
- The load average had already peaked: reading 11.08 / 38.69 / 23.59 in the right order
Build · August 15, 2026 · 1 publisher
- A 27B Apache-2.0 model in 17GB makes local inference a wiring decision, not a demo
Build · August 15, 2026 · 1 publisher
- Meta's real announcement is the split: 30B on your GPU, everything else behind the API
Build · August 14, 2026 · 6 publishers
- Qwen 3.8's Apache-licensed 27B is the one you can actually own, and its KV cache is why
Build · August 14, 2026 · 1 publisher