Skip to content

Build1 publisher3 min readPublished

Five pods green, GPU at 99 percent, queue up 70x: the Kubernetes dashboard is the wrong instrument

Pod availability, CPU and memory can all read healthy while inference degrades. The signals that matter are queue depth, time to first token and tokens per second.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Five pods green, GPU at 99 percent, queue up 70x: the Kubernetes dashboard is the wrong instrument
Generated illustration

What happened

  • The article is part of a dev.to series, following Part 2 of "AI Infrastructure for Cloud Engineers," which covered how GPUs, scheduling, autoscaling and model serving change the way AI workloads run on Kubernetes.
  • Example Kubernetes readout for an AI workload: Pods Running 5/5, CPU Usage 42%, Memory Usage 58%, Pod Restarts 0 - everything looks healthy.
  • For the same workload, the AI-specific view shows GPU utilization 99%, inference latency increasing, queue depth growing, and time to first token increasing.
  • The author concludes that in this scenario the platform is technically running but the user experience is still getting worse, and that AI observability requires connecting infrastructure health with model behavior.
  • The author frames AI observability in four layers: 1. Kubernetes infrastructure, 2. GPU / accelerator, 3. model inference, 4. end-to-end request.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A dev.to writeup in the "AI Infrastructure for Cloud Engineers" series puts two readouts from the same system side by side [1]. The Kubernetes view: 5 of 5 pods running, CPU at 42 percent, memory at 58 percent, zero pod restarts [2]. The AI view for the same workload: GPU utilization at 99 percent, inference latency increasing, queue depth growing, time to first token increasing [3]. The author's summary is the useful part: the platform is technically running while the user experience is getting worse [4].

That is the operational consequence. Pod availability and restart counts are liveness signals, and liveness is not the failure mode of a saturated inference service. Nothing crashes. Requests queue.

The post organises the problem into four layers: Kubernetes infrastructure, GPU or accelerator, model inference, and the end-to-end request [5]. The first layer still earns its keep, because when inference slows you need to rule out memory pressure, a node networking fault, or an unreachable dependency before you go looking inside the model [6]. The second layer needs utilization, GPU memory, temperature, power, device health and GPU errors, which in NVIDIA environments DCGM Exporter can publish in Prometheus-compatible form for Prometheus and Grafana [7].

The important discipline is correlation, not thresholds. According to the post, a GPU at 95 percent with a low queue, stable latency and high throughput can be perfectly healthy [8]. A GPU at 95 percent with a growing queue, rising latency and rising errors is a different system entirely [9]. Same number on the tile, opposite conclusions.

The inference layer is where the actionable metrics live: request rate, inference latency, time to first token, tokens per second, queue depth, concurrent requests, model errors and timeouts [10]. Each maps to a distinct stage of the request path, from queue through model processing to first token and response generation [11]. Inference latency is the total, time to first token is what the user waits before seeing anything, and tokens per second is generation speed after processing begins [12].

Queue depth is presented as the earliest warning. The worked example runs 2 waiting requests at 09:00, 18 at 09:05, 64 at 09:10, and 140 at 09:15 [13], a 70x increase inside fifteen minutes [14] with no crash and no restart to trigger anything [15]. That is also a scaling input: queue growth signals more inference capacity, which drains the queue [16].

The fourth layer is tracing, and the numbers make the case. An 8.2 second request tells you nothing [17]. The span breakdown attributes 40 ms to the API gateway, 70 ms to the application, 420 ms to vector search, 6.4 seconds to model inference and 950 ms to an external tool [18]. Model inference is roughly 78 percent of the wall clock [20]; the spans sum to 7.88 seconds [19], leaving about 320 ms unattributed [21], which is its own small lesson about coverage gaps in a trace.

What to watch: whether your autoscaler reads queue depth or CPU, because the second one will not move during the failure described here [3][16]; whether GPU telemetry is actually wired into the same store as pod metrics, since DCGM output is only useful if it can be correlated with the rest [7]; and whether anyone has an alert on time to first token, the one number a user directly perceives [12].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories