Skip to content

Topic

LLM Inference Serving

Field covering engines (vLLM, SGLang, Ollama) and techniques like KV caching, PagedAttention, and batching that run trained LLMs efficiently in production.

Current stories

build1 publisher

Patterned PEFT LoRA adapters run at the wrong scale in vLLM 0.30.0

vLLM 0.30.0 ignores the per-module rank and alpha patterns in PEFT LoRA adapters, and one test put the error against PEFT at 70 times the unpatterned baseline. The adapter loads without complaint, so only a config check or a side-by-side against PEFT will catch it.

Publishers:dev.to

Reality

Evidence64
Adoption
Insufficient
Hype gap+12
Incentives20
Confidence66
build1 publisher

Repacked 4-bit embeddings lift Gemma 4 decode up to 1.39x on a single L4

Repacking Gemma 4's QAT weights with 4-bit embedding tables made decode up to 1.39x faster on one SageMaker L4, according to a dev.to benchmark series. For teams serving Gemma 4 on vLLM, how the weights are stored becomes a setting to measure alongside model size.

Publishers:dev.to

Reality

Evidence55
Adoption
Insufficient
Hype gap+10
Incentives
Insufficient
Confidence50

Earlier coverage

  1. HyperPod's inference gateway scores KV cache and LoRA residency before it picks a pod

    Build · September 18, 2026 · 1 publisher

  2. Pointing the Ray head and workers at vLLM's own image removes the numpy 2.0 ABI crash

    Build · September 17, 2026 · 1 publisher

  3. An hour-long rollout and a minutes-long training step pushed Periodic Labs onto two GPU pools

    Build · September 16, 2026 · 1 publisher

  4. A 30B model that activates 3B still has to keep all 30B in VRAM

    Build · September 15, 2026 · 1 publisher

  5. Routing on the prompt's first tokens cut AWS's median time to first token by up to 77 percent

    Build · September 10, 2026 · 1 publisher

  6. NVIDIA's 2.5x concurrency figure rests on a 64K prompt reused 76 percent of the time

    Build · September 10, 2026 · 1 publisher

  7. A vLLM request crosses two IPC queues before its result reaches the caller

    Build · September 7, 2026 · 1 publisher

  8. A twelvefold longer prompt costs this RTX 3090 only 12 percent of its decode throughput

    Build · September 6, 2026 · 1 publisher

  9. One API key per platform puts every customer's system prompt in the same cache namespace

    Build · September 5, 2026 · 1 publisher

  10. Decode drags the entire model out of VRAM once per word

    Build · August 31, 2026 · 1 publisher

  11. Intel puts its Arc GPU operating knowledge inside the coding agent already installed

    Build · August 28, 2026 · 1 publisher

  12. Shadow engines cut LLM restart from 283 seconds to 7.3, and change what headroom is for

    Build · August 25, 2026 · 1 publisher

  13. AWS moves KubeRay chores into HyperPod, and the build-vs-buy math with them

    Build · August 24, 2026 · 1 publisher

  14. Ollama, vLLM, SGLang: the throughput ceiling is set by the queue, not the weights

    Build · August 22, 2026 · 1 publisher

  15. SGLang's one-GPU Qwen3.8-27B recipe is the useful half of the release

    Build · August 21, 2026 · 1 publisher

  16. The sparse-model bill arrives at serving time, and it is paid in collectives

    Build · August 21, 2026 · 1 publisher