Skip to content

Topic

Quantization formats

Reduced-precision formats (e.g., FP8, GGUF, INT4) for model weights and KV caches that shrink memory use and speed inference at some cost to fidelity.

Current stories

build1 publisher

Lokutor's Oído transcribes any English sentence on an emulated $5 ESP32-S3

Lokutor released Oído, open-source transcription of any English sentence on a $5 ESP32-S3, with every figure so far computed off the board. Whether hardware built for command lists can take open-ended speech now depends on how the real boards measure.

Publishers:dev.to

Reality

Evidence35
Adoption
Insufficient
Hype gap+30
Incentives75
Confidence40

Earlier coverage

  1. Chip Huyen puts inference at 10 to 100 times a model's training compute

    Build · September 13, 2026 · 1 publisher

  2. An unquantized model on the main thread cost a health app its screening feature

    Build · September 13, 2026 · 1 publisher

  3. Edge0's 35B would need 22 GB/s of SSD bandwidth without the file cache

    Build · September 11, 2026 · 2 publishers

  4. Bartowski broke tensors one at a time to find where GGUF bits belong

    Build · September 11, 2026 · 1 publisher

  5. 16 KB of per-row scales recover 2.2 of the 2.5 points int4 costs an embedding table

    Build · September 10, 2026 · 1 publisher

  6. NVFP4 squeezes Qwen3.8's 2.4 trillion weights onto eight B300s at 150 GB a GPU

    Build · September 9, 2026 · 1 publisher

  7. llama.cpp takes roughly half an hour to reach first token on an RTX 5090

    Build · September 8, 2026 · 1 publisher

  8. Size the model to the RAM you own before the 45-minute download

    Build · September 6, 2026 · 1 publisher

  9. One registry entry gates 16 of FreeToolHub's 21 in-browser tools on a single model

    Build · September 5, 2026 · 1 publisher

  10. Before you buy another GPU, check num_ctx and the rope base

    Build · August 22, 2026 · 1 publisher

  11. Unsloth's 10% quant claim is really about which machines can run a 27B model

    Build · August 19, 2026 · 1 publisher

  12. NVIDIA's 4-bit Nemotron shows what aggressive quantization costs: you have to retrain for it

    Build · August 17, 2026 · 1 publisher

  13. Your 2026 GPU Decision Is Arithmetic: Bytes Per Parameter, Times Parameters, Plus Cache

    Build · August 16, 2026 · 1 publisher

  14. A 27B Apache-2.0 model in 17GB makes local inference a wiring decision, not a demo

    Build · August 15, 2026 · 1 publisher