Skip to content

project

llama.cpp

Open-source C/C++ library for running LLM inference locally on CPUs and GPUs, using the GGUF model format; includes tools like llama-server and llama-cli.

Known aliases

  • ggml-org/llama.cpp
  • llama-bench
  • llama-cli
  • llama cpp
  • llama_cpp
  • llama.cpp
  • llama.cpp reference
  • llama-perplexity
  • llama-quant.cpp
  • llama-server

Relationships

No evidence-backed relationships are recorded.

Current stories

build3 publishers

Typesafe's Jev API gains an open-weight rival in Cloudflare's Clef decision models

Cloudflare released two Jev-API-compatible decision models, Clef and Clef-flash, on Workers AI and as Apache 2.0 weights on Hugging Face. Typed classification steps in agent code can now move between providers or onto owned hardware, as long as they stay inside the text-only, 32k-context features Jev supports.

Perspective Coverage

3 publishers
Builder
Builder 53%
Operator
Operator 27%
Investor
Investor 20%

Reality

Evidence50
Adoption18
Hype gap+35
Incentives70
Confidence60
build3 publishers

llama.cpp brings TypeSafe's typed-decision format to five open model families on local hardware

llama.cpp merged a /v1/systemone endpoint on October 2 that returns typed answers with probabilities from five open model families running locally. Whether it can replace hosted classification calls depends on calibration that each team has to measure on its own data.

Perspective Coverage

3 publishers
Builder
Builder 48%
Operator
Operator 29%
Investor
Investor 23%

Reality

Evidence50
Adoption15
Hype gap+5
Incentives55
Confidence55

Earlier coverage

  1. NVFP4 and cache reuse cut MLPerf's edge agent workload to 24 minutes on one Jetson Thor

    Build · September 16, 2026 · 1 publisher

  2. A 4 GB laptop GPU decodes quantised Gemma 4 at 4.27x the CPU rate on 1598 MiB

    Build · September 16, 2026 · 1 publisher

  3. Filling Qwen 3.8 27B's native context costs about as much memory as its weights

    Build · September 15, 2026 · 1 publisher

  4. Running llama-server puts context, KV cache and GPU placement in your command line

    Build · September 14, 2026 · 1 publisher

  5. A diffusion drafter lost to Gemma's own Assistant model on a 12GB RTX 3060

    Build · September 14, 2026 · 1 publisher

  6. A million serverless briefings for $48 implies a Lambda rate of $0.000004 per GB-second

    Build · September 12, 2026 · 1 publisher

  7. A backend that detects the AMD GPU can still leave operations on the CPU

    Build · September 12, 2026 · 1 publisher

  8. 6,935 exposed Ollama servers answered an internet scan without asking for credentials

    Security · September 11, 2026 · 1 publisher

  9. Bartowski broke tensors one at a time to find where GGUF bits belong

    Build · September 11, 2026 · 1 publisher

  10. Overnight laptop runs took over most of one Rust developer's Opus coding work

    Build · September 10, 2026 · 1 publisher

  11. A 6% driver reserve decides which models fit on a $2,000 pair of P40s

    Build · September 8, 2026 · 1 publisher

  12. llama.cpp takes roughly half an hour to reach first token on an RTX 5090

    Build · September 8, 2026 · 1 publisher

  13. Size the model to the RAM you own before the 45-minute download

    Build · September 6, 2026 · 1 publisher

  14. NVIDIA's edge-agent case rests on compact open models matching data-center capabilities

    Build · September 4, 2026 · 1 publisher

  15. RamaLama ships models as OCI images you can inspect and sign

    Build · August 31, 2026 · 1 publisher

  16. A 7B model at 11 tokens per second cleared the bar for private contract Q&A

    Build · August 30, 2026 · 1 publisher

  17. Reading one ambiguous config key correctly pushed a passing cost model to 3.4% error

    Build · August 30, 2026 · 1 publisher

  18. Running the model on the laptop turns a subscription line into a maintenance chore

    Product · August 29, 2026 · 1 publisher

  19. Four-bit weights leave 6 GB on a 24 GB card for KV cache and vision tensors

    Build · August 29, 2026 · 1 publisher

  20. Intel puts its Arc GPU operating knowledge inside the coding agent already installed

    Build · August 28, 2026 · 1 publisher

  21. Prefill ate 85% of a 291-second answer, and the fix was a dedup key and a cache slot

    Build · August 26, 2026 · 1 publisher

  22. Mac Studio M5 Ultra vs DGX Spark: capacity says what fits, bandwidth says what you wait for

    Build · August 25, 2026 · 1 publisher

  23. Parallels ships OpenGL 4.3 on Apple silicon, and the fleet question moves to which chip you own

    Product · August 25, 2026 · 1 publisher

  24. Pi ships four tools and no sandbox, so the guardrails are on your build list

    Build · August 23, 2026 · 1 publisher

  25. A 284B model at 25 tokens a second on one 5090, and 192 GiB of DDR5 doing the quiet part

    Build · August 22, 2026 · 1 publisher

  26. Base Compute hands kernel tuning to agents; the carryover claim is the unmeasured part

    Build · August 21, 2026 · 1 publisher

  27. The flash_attn error in llama.cpp is a layout constraint, and it decides your context window

    Build · August 21, 2026 · 1 publisher

  28. Unsloth's 10% quant claim is really about which machines can run a 27B model

    Build · August 19, 2026 · 1 publisher

  29. Inco AI's DFlash 2: 21% longer accepted drafts for 1.3% latency and 18.5M parameters

    Build · August 19, 2026 · 1 publisher

  30. You Procured Qwen. Your Edge Boxes Are Running Somebody Else's File.

    Leadership · August 18, 2026 · 1 publisher

  31. A refusal-stripped 27B model now ships as a 17.9 GB llama.cpp pull

    Build · August 16, 2026 · 1 publisher

  32. The load average had already peaked: reading 11.08 / 38.69 / 23.59 in the right order

    Build · August 15, 2026 · 1 publisher

  33. A 27B Apache-2.0 model in 17GB makes local inference a wiring decision, not a demo

    Build · August 15, 2026 · 1 publisher

  34. Meta's real announcement is the split: 30B on your GPU, everything else behind the API

    Build · August 14, 2026 · 6 publishers

  35. Qwen 3.8's Apache-licensed 27B is the one you can actually own, and its KV cache is why

    Build · August 14, 2026 · 1 publisher