Skip to content

standard

GGUF

GGUF is a binary file format for storing quantized large language model weights, used by llama.cpp and compatible tools to run models on consumer hardware.

Known aliases

  • 4-bit GGUF
  • GGML Universal File
  • .gguf
  • GGUF format
  • Q4 GGUF

Relationships

No evidence-backed relationships are recorded.

Current stories

build1 publisher

Size the model to the RAM you own before the 45-minute download

A 15M-parameter model streams English text on a 2007 PSP at about one token per second. That is the extreme end of a sizing rule. The harder half of that rule is checking whether the file that fits is a format its own maintainer recommends.

Publishers:dev.to

Reality

Evidence38
Adoption31
Hype gap+12
Incentives58
Confidence46
build1 publisher

RamaLama ships models as OCI images you can inspect and sign

Treating the runtime and the weights as a container image buys you a hardened default docker run and a signable artifact. On Apple Silicon, though, that same boundary costs you the GPU. One hands-on run puts a number on that cost.

Publishers:dev.to

Reality

Evidence52
Adoption20
Hype gap−5
Incentives30
Confidence48

Earlier coverage

  1. Binding a local model server to 0.0.0.0 hands the LAN an unauthenticated API

    Build · August 30, 2026 · 1 publisher

  2. Ollama, vLLM, SGLang: the throughput ceiling is set by the queue, not the weights

    Build · August 22, 2026 · 1 publisher

  3. "Local" Is A Statement About Inference, Not About Sockets

    Build · August 20, 2026 · 1 publisher

  4. Unsloth's 10% quant claim is really about which machines can run a 27B model

    Build · August 19, 2026 · 1 publisher

  5. Ornith-1.0's benchmarks are fine. Ollama can't parse its tool calls.

    Build · August 18, 2026 · 1 publisher

  6. You Procured Qwen. Your Edge Boxes Are Running Somebody Else's File.

    Leadership · August 18, 2026 · 1 publisher

  7. A refusal-stripped 27B model now ships as a 17.9 GB llama.cpp pull

    Build · August 16, 2026 · 1 publisher

  8. Your 2026 GPU Decision Is Arithmetic: Bytes Per Parameter, Times Parameters, Plus Cache

    Build · August 16, 2026 · 1 publisher

  9. A .keras config can carry a marshalled Python code object, and load_model runs it

    Build · August 16, 2026 · 1 publisher

  10. A 27B Apache-2.0 model in 17GB makes local inference a wiring decision, not a demo

    Build · August 15, 2026 · 1 publisher