Skip to content

project

vLLM

Open-source library for fast, memory-efficient LLM inference and serving, supporting diverse hardware like GPUs, TPUs, and Gaudi.

Known aliases

  • vLLM
  • vLLM 0.22
  • vLLM 0.25.1
  • vLLM 0.27.1
  • vllm bench serve
  • vLLM Deep Learning Container
  • vLLM-Omni
  • vllm-openai
  • vLLM project
  • vllm-project/vllm
  • vllm serve
  • vLLM-TPU
  • vLLM V1
  • vllm/vllm-openai
  • vLLM-XPU

Relationships

No evidence-backed relationships are recorded.

Current stories

build4 publishers

Aleph Alpha's open Kolibri model routes each token through 3.46B of its 78.1B parameters

Aleph Alpha released Kolibri, an Apache 2.0 German-English model that activates 3.46B of its 78.1B parameters per token. Each token costs about as much compute as a small model, yet a team hosting it in Europe still has to fit every expert in memory.

Perspective Coverage

4 publishers
Builder
Builder 47%
Operator
Operator 36%
Investor
Investor 17%

Reality

Evidence62
Adoption
Insufficient
Hype gap+15
Incentives68
Confidence66
build1 publisher

Moonshot publishes open weights for its trillion-parameter K2.7-Code model

Moonshot AI released open weights for Kimi K2.7-Code, a trillion-parameter coding model that activates 32 billion parameters per token. Its headline gains come from Moonshot's own benchmarks, so teams paying for proprietary agents have to measure it on their own code.

Publishers:dev.to

Reality

Evidence35
Adoption
Insufficient
Hype gap+20
Incentives55
Confidence40
invest3 publishers

Nebius buys idle-GPU startup Inferize for an estimated $100 million to $150 million

Nebius bought Inferize, a 10-month-old Israeli startup with 17 staff, for an estimated $100 million to $150 million, according to Calcalist. The money goes on software that keeps inference GPUs from sitting idle, and Nebius is still adding data center capacity alongside it.

Publishers:calcalistech.comnebius.comruntimewire.com

Perspective Coverage

3 publishers
Builder
Builder 38%
Operator
Operator 32%
Investor
Investor 30%

Reality

Evidence55
Adoption12
Hype gap+30
Incentives70
Confidence60
build1 publisher

Patterned PEFT LoRA adapters run at the wrong scale in vLLM 0.30.0

vLLM 0.30.0 ignores the per-module rank and alpha patterns in PEFT LoRA adapters, and one test put the error against PEFT at 70 times the unpatterned baseline. The adapter loads without complaint, so only a config check or a side-by-side against PEFT will catch it.

Publishers:dev.to

Reality

Evidence64
Adoption
Insufficient
Hype gap+12
Incentives20
Confidence66
build1 publisher

Repacked 4-bit embeddings lift Gemma 4 decode up to 1.39x on a single L4

Repacking Gemma 4's QAT weights with 4-bit embedding tables made decode up to 1.39x faster on one SageMaker L4, according to a dev.to benchmark series. For teams serving Gemma 4 on vLLM, how the weights are stored becomes a setting to measure alongside model size.

Publishers:dev.to

Reality

Evidence55
Adoption
Insufficient
Hype gap+10
Incentives
Insufficient
Confidence50
build3 publishers

Granite 4.2 ships a self-hostable reasoning tier under Apache 2.0, and its data says coding agent

IBM's 3B, 8B and 30B dense models all get a thinking switch and native tool calling, but only the two larger ones get agentic RL, and the tuning mixture leans hard on software engineering.

Perspective Coverage

3 publishers
Builder
Builder 63%
Operator
Operator 28%
Investor
Investor 9%

Reality

Evidence55
Adoption
Insufficient
Hype gap+25
Incentives60
Confidence72

Earlier coverage

  1. HyperPod holds the inference deployment until every target node has cached the weights

    Build · September 10, 2026 · 1 publisher

  2. Kimi Linear's 75% KV cache cut comes from making three of every four layers recurrent

    Build · September 25, 2026 · 1 publisher

  3. rocminfo reports one MI300X framebuffer three times, once per global memory pool

    Build · September 16, 2026 · 1 publisher

  4. A 4-bit Gemma 4 26B on one L4 trails TypeSafe's Jev by 2.1 points overall

    Build · September 23, 2026 · 1 publisher

  5. Folding 92 layers into one Pallas call put Kimi K3 at 709 tokens a second on TPU v7

    Build · September 23, 2026 · 1 publisher

  6. Kaitchup traces Bonsai 2's 98.2% retention figure to unpacked weights on an H100

    Build · September 22, 2026 · 1 publisher

  7. vLLM measured its portability layer at 3.4 percent below native throughput on an H100

    Build · September 22, 2026 · 1 publisher

  8. A CreateAIBenchmarkJob call ramps a SageMaker endpoint from 64 to 1,024 concurrent requests

    Build · September 22, 2026 · 1 publisher

  9. Crusoe's Series F values a $140 billion backlog at 22 cents on the dollar

    Invest · September 22, 2026 · 1 publisher

  10. Shopify's continual learning loop cuts serving costs 96%; case study separately details judges calibrated from 25 annotated samples

    Build · September 22, 2026 · 1 publisher

  11. VoxCPM2 trades the per-character speech bill for a GPU and a base-URL change

    Build · September 21, 2026 · 1 publisher

  12. White Circle publishes the harness behind Halo's 2.3x TRL throughput claim

    Build · September 21, 2026 · 1 publisher

  13. Paging the KV cache into 16-token blocks caps the waste at 960 KB per sequence

    Build · September 21, 2026 · 1 publisher

  14. A resident Whisper model plus embedding model together burned 2 euros of GPU electricity across 30 days

    Build · September 20, 2026 · 1 publisher

  15. Crusoe's $3.9bn round prices a contract book that runs five gigawatts ahead of delivery

    Invest · September 19, 2026 · 2 publishers

  16. A local proxy convinces the ChatGPT desktop app it is still talking to OpenAI

    Build · September 19, 2026 · 1 publisher

  17. SageMaker's fallback instance can serve a quantized copy of the same model

    Build · September 18, 2026 · 1 publisher

  18. A tile-size clamp in vLLM's Triton kernel gets Gemma 4 running on a Tesla T4

    Build · September 18, 2026 · 1 publisher

  19. NVIDIA's AIPerf drops the Perf Analyzer stack for worker processes over ZMQ

    Build · September 18, 2026 · 1 publisher

  20. Kiro and Claude Code both picked a TGI container that could not load Qwen3

    Build · September 18, 2026 · 1 publisher

  21. Vals put Hy4 Preview first among open-weight models on code migration at $3.41 a test

    Build · September 17, 2026 · 1 publisher

  22. llama.cpp's -ngl flag keeps a 9B model on a 6GB card by leaving 28 layers on the CPU

    Build · September 17, 2026 · 1 publisher

  23. Pointing the Ray head and workers at vLLM's own image removes the numpy 2.0 ABI crash

    Build · September 17, 2026 · 1 publisher

  24. Thought-anchor scores change when a second model resamples the same trace

    Build · September 16, 2026 · 1 publisher

  25. NVIDIA's 3.7x Vera Rubin figure comes from Qwen3-VL on vLLM and Dynamo

    Build · September 16, 2026 · 1 publisher

  26. One slash in a Host header moves the path Starlette's middleware checks

    Build · September 16, 2026 · 1 publisher

  27. Red Hat clocks sandbox isolation at under 5 percent of an agent request's latency

    Product · September 15, 2026 · 1 publisher

  28. Whoever implements the server half of Responses picks your retrieval and veto defaults

    Build · September 15, 2026 · 1 publisher

  29. Rubin shows 7x tokens per megawatt and over 2x profit per gigawatt versus Blackwell, SemiAnalysis says

    Invest · September 14, 2026 · 1 publisher

  30. Red Hat traces the inference bill to 140GB of memory reads per token

    Product · September 14, 2026 · 1 publisher

  31. Mixing self-hosted Qwen with Bedrock Claude costs you telemetry, not a rewrite

    Build · August 14, 2026 · 1 publisher

  32. A backend that detects the AMD GPU can still leave operations on the CPU

    Build · September 12, 2026 · 1 publisher

  33. 6,935 exposed Ollama servers answered an internet scan without asking for credentials

    Security · September 11, 2026 · 1 publisher

  34. Renaming FastMCP to MCPServer closes stdio before an unpinned server finishes its handshake

    Build · September 11, 2026 · 1 publisher

  35. Overnight laptop runs took over most of one Rust developer's Opus coding work

    Build · September 10, 2026 · 1 publisher

  36. Routing on the prompt's first tokens cut AWS's median time to first token by up to 77 percent

    Build · September 10, 2026 · 1 publisher

  37. NVIDIA's 2.5x concurrency figure rests on a 64K prompt reused 76 percent of the time

    Build · September 10, 2026 · 1 publisher

  38. Why LLM output is hard to reproduce: it's not just concurrency and floating point

    Science · September 10, 2026 · 1 publisher

  39. NVFP4 squeezes Qwen3.8's 2.4 trillion weights onto eight B300s at 150 GB a GPU

    Build · September 9, 2026 · 1 publisher

  40. Judge model choice swings AI-Infra-Guard's false positive rate fifteenfold

    Security · September 9, 2026 · 1 publisher