Skip to content

framework

SGLang

SGLang is an open-source inference-serving framework for large language models, optimized for high-throughput deployment via efficient KV-cache reuse.

Known aliases

  • SGLang
  • sglang.launch_server
  • SGLang project
  • SGLang Spec V2
  • SGLang-XPU

Relationships

No evidence-backed relationships are recorded.

Current stories

build1 publisher

Moonshot publishes open weights for its trillion-parameter K2.7-Code model

Moonshot AI released open weights for Kimi K2.7-Code, a trillion-parameter coding model that activates 32 billion parameters per token. Its headline gains come from Moonshot's own benchmarks, so teams paying for proprietary agents have to measure it on their own code.

Publishers:dev.to

Reality

Evidence35
Adoption
Insufficient
Hype gap+20
Incentives55
Confidence40
build3 publishers

Granite 4.2 ships a self-hostable reasoning tier under Apache 2.0, and its data says coding agent

IBM's 3B, 8B and 30B dense models all get a thinking switch and native tool calling, but only the two larger ones get agentic RL, and the tuning mixture leans hard on software engineering.

Perspective Coverage

3 publishers
Builder
Builder 63%
Operator
Operator 28%
Investor
Investor 9%

Reality

Evidence55
Adoption
Insufficient
Hype gap+25
Incentives60
Confidence72
build5 publishers

Alibaba ships the Qwen4 architecture as open weights before the flagship exists

Qwen3.8-Flash-Next puts 36 Gated DeltaNet layers and 12 sparse-attention layers on Hugging Face, which means the retrieval budget Qwen4 will inherit is something you can measure against your own traces now.

Perspective Coverage

5 publishers
Builder
Builder 52%
Operator
Operator 28%
Investor
Investor 20%

Reality

Evidence58
Adoption52
Hype gap+32
Incentives76
Confidence71

Earlier coverage

  1. GLM-5.3-Flash benchmarks its tenth-of-the-price claim against its own predecessor

    Leadership · September 5, 2026 · 1 publisher

  2. Peak-hour pricing pushes DeepSeek's new vision model past Gemini on the invoice test

    Build · August 31, 2026 · 2 publishers

  3. DeepSeek's MIT-licensed V4-Pro hands API buyers a credible walk-away option

    Leadership · August 30, 2026 · 1 publisher

  4. Four-bit weights leave 6 GB on a 24 GB card for KV cache and vision tensors

    Build · August 29, 2026 · 1 publisher

  5. Intel puts its Arc GPU operating knowledge inside the coding agent already installed

    Build · August 28, 2026 · 1 publisher

  6. Streaming tool-call deltas turn a base-URL swap into a per-model parser project

    Build · August 28, 2026 · 1 publisher

  7. SPEED-Bench re-tests speculative decoding at the batch size you actually serve

    Build · August 27, 2026 · 1 publisher

  8. Z.ai's cost-parity claim on Chinese accelerators rests on model design as much as silicon

    Leadership · August 27, 2026 · 1 publisher

  9. Two buyers, one price: Patel says 2027's new compute is already half spoken for

    Build · August 25, 2026 · 1 publisher

  10. Shadow engines cut LLM restart from 283 seconds to 7.3, and change what headroom is for

    Build · August 25, 2026 · 1 publisher

  11. Nvidia's agent-workload lead scales with interactivity, and the second-source budget line does not

    Leadership · August 25, 2026 · 1 publisher

  12. MiniMax maps H3 from one 24GB card to SGLang, and keeps the interpreter in-house

    Build · August 24, 2026 · 1 publisher

  13. A benchmark that replays real agent sessions gives back less of the generational win

    Build · August 24, 2026 · 1 publisher

  14. A 27B model reportedly beat a license check in 30 minutes. Nobody has seen the binary.

    Build · August 23, 2026 · 1 publisher

  15. Ollama, vLLM, SGLang: the throughput ceiling is set by the queue, not the weights

    Build · August 22, 2026 · 1 publisher

  16. SGLang's one-GPU Qwen3.8-27B recipe is the useful half of the release

    Build · August 21, 2026 · 1 publisher

  17. The real disclosure in Qwen3.8-Max is the rack: 2.4T open weights, 72 GPUs, 4K tokens/sec

    Science · August 20, 2026 · 1 publisher

  18. Inco AI's DFlash 2: 21% longer accepted drafts for 1.3% latency and 18.5M parameters

    Build · August 19, 2026 · 1 publisher

  19. Your vLLM Manifest Would Boot SGLang Too, And That Is the Problem

    Build · August 18, 2026 · 1 publisher

  20. Dual 3090s, no NVLink: the serving stack broke long before the model did

    Build · August 18, 2026 · 1 publisher

  21. Qwen3.8's 27B dense checkpoint is the one operators can actually host

    Build · August 14, 2026 · 1 publisher