Skip to content

Build1 publisher3 min readPublished

Your vLLM Manifest Would Boot SGLang Too, And That Is the Problem

A Kubernetes serving guide makes a point most teams skip: the YAML you reviewed is identical across five engines, and everything it hides is what breaks in production.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • In Part 6 of the series, the author deployed Qwen/Qwen2.5-1.5B-Instruct on a Kubernetes GPU node and called it with curl; it answered because the pod ran a serving engine called vLLM.
  • vLLM quietly did the hard parts: downloaded weights, loaded them onto the GPU, started an OpenAI-compatible HTTP server, and handled token generation.
  • SGLang, TGI, Triton and TensorRT-LLM can all sit in the same container and serve the same model behind the same Kubernetes Service, and the kubectl apply looks almost identical.
  • What changes between engines is everything that matters for production: startup time, memory behavior, latency, throughput, batching, observability, and how much wiring you have to do yourself.
  • Kubernetes provides pods, scheduling, GPU allocation via the device plugin, Secrets, Services, health checks, rolling updates and networking; all of it is real and necessary, and none of it generates tokens.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A Kubernetes serving guide published on dev.to makes a point worth repeating to any team that shipped its first LLM API and moved on: the pod that answered your curl request worked because vLLM was inside the container, quietly downloading weights, loading them onto the GPU, starting an OpenAI-compatible HTTP server and generating tokens [1][2]. According to the same piece, SGLang, TGI, Triton and TensorRT-LLM could sit in that same container and serve the same model behind the same Kubernetes Service, and the `kubectl apply` would look almost identical [3].

That symmetry is the trap. The manifest is the artifact your team reviews, and it is precisely the artifact that does not encode the decision. What changes between engines, per the author, is startup time, memory behavior, latency, throughput, batching, observability, and how much wiring you have to do yourself [4].

The division of labor is worth stating flatly, because it is where the misdiagnosis starts. Kubernetes gives you pods, scheduling, GPU allocation through the device plugin, Secrets, Services, health checks, rolling updates and networking, and none of that generates tokens [5]. The engine loads weights, initializes the tokenizer and runtime, manages the KV cache as requests arrive, batches to keep the GPU busy, runs the forward passes and exposes metrics [6]. So a pod sitting in `Running` with healthy CPU and memory can still have a broken batching strategy, a KV cache that is too small, or a queue backing up silently, and Kubernetes will not report it because Kubernetes does not know [7]. When the API is slow, the guide's advice is to look at the engine first [8].

None of this is an argument against vLLM. It is open source, broadly supported across hardware, ships an OpenAI-compatible API out of the box, and implements PagedAttention for memory-efficient KV cache management [10]. The serve command is one line [11]. The author's honest caveat is narrower than the usual benchmark fight: vLLM is less obvious for heavily structured generation, complex multi-step prompt programs, and workloads needing fine control over prefix caching across agentic chains, and while vLLM has answers for all of these, other engines were built around some of them first [12].

The concrete alternative in the material is SGLang, which the author says has been taking vLLM's territory over the past year [13]. Its RadixAttention prefix cache detects when requests share a common prompt prefix and reuses the computed work, which the author frames as a latency and cost win for agentic workloads, RAG and repeated prompt scaffolding [14]. SGLang also carries structured generation, multi-LoRA support, early prefill-decode disaggregation and OpenAI API compatibility, and the author names xAI among its large production deployments [15]. The caveat is the useful part: RadixAttention helps when the beginning token sequence is literally the same, not when prompts are semantically similar [16].

What to watch before you change engines: measure whether your traffic actually shares leading tokens, because that single property decides whether a prefix cache pays for itself [16]. Instrument the engine's own metrics rather than pod status, since queue depth and cache pressure live there [6][7]. And treat startup time and memory behavior as rollout properties you test, not footnotes [4]. The supplied guide details vLLM and SGLang and stops short of TGI, Triton and TensorRT-LLM, so those three remain names on a list rather than a comparison [17].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories