Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished

Probes cannot fix bad addressing: what a browser-run Kubernetes cluster actually shows

An ngrok post ports more than 100,000 lines of Kubernetes Go to TypeScript to demonstrate restart loops and dropped requests. The instructive case is the one a correct probe does not prevent.

The Engineer · Build desk

How we use AISend a correction

What happened

  • An ngrok post teaches Kubernetes probes through demos running on webernetes, a partial TypeScript port carrying over 100,000 lines of translated Kubernetes Go code.
  • The author says he checked the simulated behaviour against k3s and found a bug in Kubernetes along the way.
  • With no probe set, a container is counted Ready the instant it starts, including while it is still initialising and not yet listening on port 8080.
  • CrashLoopBackOff defaults to a 10 second delay that doubles per crash to a 5 minute maximum; the demo compresses it to 3 seconds.

Why it matters

  • constraint Readiness only withdraws a pod from endpoint lists, so any caller holding a pod IP keeps hitting it while it is unready; the fix belongs on the client's addressing, not in the probe block.
  • cost Once the backoff ceiling is reached, the operator is paying 5 minutes of dead pod per attempt, which at one request every 2 seconds is 150 failures per gap before anyone reads the alert.
  • decision periodSeconds times failureThreshold is a deadline the kubelet enforces by killing the container, so it has to be sized against measured startup time rather than copied from an example.
  • capability If the browser port really tracks k3s, probe combinations can be exercised without standing up a cluster, which lowers the cost of testing a rollout theory before it reaches production.

The default that produces the outage is the absence of a probe, not the presence of a bad one. With nothing configured, Kubernetes treats a container as Ready the moment it starts, including the seconds it spends initialising before it listens on port 8080 [6]. The container is up, it is advertised as healthy, and requests arriving in that window fail [6]. Nothing malfunctioned. The kubelet did what it was told, and it was told nothing.

The cost lives in the restart curve. CrashLoopBackOff starts at 10 seconds and doubles with each crash up to a ceiling of 5 minutes [7]. Run the sequence: 10, 20, 40, 80, 160, then 320 clipped to 300, so the sixth backoff already sits at the cap and about 10 minutes of wall clock is gone [14]. After that the container gets 12 start attempts an hour [15]. If the caller behaves like the post's pod-b, one request every 2 seconds [12], each capped gap is 150 requests with nowhere to land [16]. That is the mechanism behind "restart loops that take hours to recover from" [4]: a doubling that outruns the person watching it.

The more useful demo is the one where the probe is correct and the requests still fail. Add an httpGet startup probe and the pod reports itself not ready until the first GET /startup returns a status in the 200-399 range [9][11]. The client in the demo keeps failing anyway, because it is pointed at pod-a's IP address directly, which never consults readiness [2]. Readiness gates endpoints, not packets. It cannot correct a caller that already knows the address, and in a real cluster that caller is an ingress controller, a load balancer, or another service holding a stale target [12].

The probe budget is the other misconfiguration with teeth: periodSeconds multiplied by failureThreshold, roughly 5 seconds in the post's example, after which the kubelet kills the container rather than waiting longer [9]. Set that budget below a container's honest startup time and slow becomes dead, and dead re-enters the backoff schedule above [17]. Each node runs its own kubelet making these decisions locally [10], so the failure is per-node and looks like flapping rather than a policy error.

On the tooling itself, the author reports porting more than 100,000 lines of Kubernetes Go to TypeScript so the simulated cluster runs in the browser, verifying demo behaviour against k3s, and finding a bug in Kubernetes in the process [5][1]. The bug is announced and deferred to a later section [3]. Until it is named, the supportable claim is the narrower one: the port tracks k3s closely enough that its author trusted a divergence rather than assuming his own translation was wrong. That earns it standing as a teaching rig. It is not yet evidence about upstream. Also worth noting for anyone reading the demos literally: the NotReady label is the author's shorthand, since the API exposes a Ready condition that is True, False, or Unknown [13], and the backoff was compressed from 10 seconds to 3 for the demo [7], which makes the visible loop kinder than the one in production.

What to watch

  • Whether the promised later section names the Kubernetes bug and links a filed issue with a reproducer.
  • Whether the multi-replica section shows readiness actually removing a pod from Service endpoints, closing the gap the pod-IP demo leaves open.
  • Whether webernetes is published in a form others can run and diff against a real kubelet.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence54
Adoption
Insufficient
Hype gap+14
Incentives58
Confidence48
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    The author says he verified the behaviour of the demos against k3s and found a bug in Kubernetes.

  2. [2]

    Even when pod-a is not ready, pod-b's requests still fail during startup because pod-b is configured to send requests directly to pod-a's IP address, bypassing the readiness mechanism.

  3. [3]

    The post announces the Kubernetes bug with 'More on that later' and does not identify or describe it in the section that sets up the probe demos.

Sources

1 independent publisher whose own reporting we read for this story.

  1. ngrok.com

    1 article · August 22, 2026

    How Kubernetes probes work

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

  • Container Restart and Backoff BehaviourFollow
  • Interactive Developer EducationFollow
  • Service Discovery and AddressingFollow
  • Kubernetes Health ProbesFollow
Loading related stories