BuildNot yet confirmed elsewhere1 publisher3 min readPublished
Probes cannot fix bad addressing: what a browser-run Kubernetes cluster actually shows
An ngrok post ports more than 100,000 lines of Kubernetes Go to TypeScript to demonstrate restart loops and dropped requests. The instructive case is the one a correct probe does not prevent.
The Engineer · Build desk
What happened
- An ngrok post teaches Kubernetes probes through demos running on webernetes, a partial TypeScript port carrying over 100,000 lines of translated Kubernetes Go code.
- The author says he checked the simulated behaviour against k3s and found a bug in Kubernetes along the way.
- With no probe set, a container is counted Ready the instant it starts, including while it is still initialising and not yet listening on port 8080.
- CrashLoopBackOff defaults to a 10 second delay that doubles per crash to a 5 minute maximum; the demo compresses it to 3 seconds.
Why it matters
- constraint Readiness only withdraws a pod from endpoint lists, so any caller holding a pod IP keeps hitting it while it is unready; the fix belongs on the client's addressing, not in the probe block.
- cost Once the backoff ceiling is reached, the operator is paying 5 minutes of dead pod per attempt, which at one request every 2 seconds is 150 failures per gap before anyone reads the alert.
- decision periodSeconds times failureThreshold is a deadline the kubelet enforces by killing the container, so it has to be sized against measured startup time rather than copied from an example.
- capability If the browser port really tracks k3s, probe combinations can be exercised without standing up a cluster, which lowers the cost of testing a rollout theory before it reaches production.
The default that produces the outage is the absence of a probe, not the presence of a bad one. With nothing configured, Kubernetes treats a container as Ready the moment it starts, including the seconds it spends initialising before it listens on port 8080 [6]. The container is up, it is advertised as healthy, and requests arriving in that window fail [6]. Nothing malfunctioned. The kubelet did what it was told, and it was told nothing.
The cost lives in the restart curve. CrashLoopBackOff starts at 10 seconds and doubles with each crash up to a ceiling of 5 minutes [7]. Run the sequence: 10, 20, 40, 80, 160, then 320 clipped to 300, so the sixth backoff already sits at the cap and about 10 minutes of wall clock is gone [14]. After that the container gets 12 start attempts an hour [15]. If the caller behaves like the post's pod-b, one request every 2 seconds [12], each capped gap is 150 requests with nowhere to land [16]. That is the mechanism behind "restart loops that take hours to recover from" [4]: a doubling that outruns the person watching it.
The more useful demo is the one where the probe is correct and the requests still fail. Add an httpGet startup probe and the pod reports itself not ready until the first GET /startup returns a status in the 200-399 range [9][11]. The client in the demo keeps failing anyway, because it is pointed at pod-a's IP address directly, which never consults readiness [2]. Readiness gates endpoints, not packets. It cannot correct a caller that already knows the address, and in a real cluster that caller is an ingress controller, a load balancer, or another service holding a stale target [12].
The probe budget is the other misconfiguration with teeth: periodSeconds multiplied by failureThreshold, roughly 5 seconds in the post's example, after which the kubelet kills the container rather than waiting longer [9]. Set that budget below a container's honest startup time and slow becomes dead, and dead re-enters the backoff schedule above [17]. Each node runs its own kubelet making these decisions locally [10], so the failure is per-node and looks like flapping rather than a policy error.
On the tooling itself, the author reports porting more than 100,000 lines of Kubernetes Go to TypeScript so the simulated cluster runs in the browser, verifying demo behaviour against k3s, and finding a bug in Kubernetes in the process [5][1]. The bug is announced and deferred to a later section [3]. Until it is named, the supportable claim is the narrower one: the port tracks k3s closely enough that its author trusted a divergence rather than assuming his own translation was wrong. That earns it standing as a teaching rig. It is not yet evidence about upstream. Also worth noting for anyone reading the demos literally: the NotReady label is the author's shorthand, since the API exposes a Ready condition that is True, False, or Unknown [13], and the backoff was compressed from 10 seconds to 3 for the demo [7], which makes the visible loop kinder than the one in production.
What to watch
- Whether the promised later section names the Kubernetes bug and links a filed issue with a reproducer.
- Whether the multi-replica section shows readiness actually removing a pod from Service endpoints, closing the gap the pod-IP demo leaves open.
- Whether webernetes is published in a form others can run and diff against a real kubelet.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence54
- Adoption
- Insufficient
- Hype gap+14
- Incentives58
- Confidence48
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The author says he verified the behaviour of the demos against k3s and found a bug in Kubernetes.
ReportedSupportedSource: ngrok blog post2 sources— create a free account to open themView cited source - [2]
Even when pod-a is not ready, pod-b's requests still fail during startup because pod-b is configured to send requests directly to pod-a's IP address, bypassing the readiness mechanism.
- [3]
The post announces the Kubernetes bug with 'More on that later' and does not identify or describe it in the section that sets up the probe demos.
- [4]
The post states its aim is to show how probes make an application more resilient and prevent avoidable mistakes such as restart loops that take hours to recover from and dropping requests during rollouts.
- [5]
Every interactive demo in the post uses webernetes, the author's partial port of Kubernetes to TypeScript containing more than 100,000 lines of ported Kubernetes Go code, which runs a simulated cluster in the browser.
- [6]
With no probe configured, Kubernetes considers the container Ready as soon as it starts even though it is still doing startup work and not listening on port 8080, and requests sent during that period fail.
- [7]
After a second crash Kubernetes imposes CrashLoopBackOff; the default delay is 10 seconds, doubling with each crash up to a maximum wait of 5 minutes, and the author shortened it to 3 seconds for the demo.
- [8]
Kubernetes has three probe types: startup probes determine whether the application has started, readiness probes whether it is ready to receive traffic, and liveness probes whether it needs to be restarted.
- [9]
The example startup probe is an httpGet probe sending GET /startup on port 8080; status codes 200-399 count as success; it runs every periodSeconds and may fail failureThreshold consecutive times before Kubernetes kills the container, giving the container about 5 seconds to finish startup work.
- [10]
Probes are sent by the kubelet; each node in the cluster has its own kubelet, responsible for ensuring the right pods are running and being probed on that node.
- [11]
With a startup probe configured, the pod shows as not ready after a restart and only becomes Ready after the first startup probe succeeds; not ready is the default for pods whose containers have a startup probe.
- [12]
The demo adds pod-b, which sends a request to pod-a every 2 seconds and stands in for any source of client traffic such as an ingress controller, a load balancer, or inter-service requests.
- [13]
The author notes there is no NotReady condition in the API; a pod has a Ready condition that can be True, False, or Unknown, and NotReady was used in the demos for brevity.
- [14]
Under the default schedule the sixth backoff interval is the first at the 5-minute cap, and about 610 seconds (roughly 10 minutes) of backoff has elapsed by then.
- [15]
At the 5-minute backoff ceiling a crashing container gets 12 start attempts per hour.
- [16]
A client sending one request every 2 seconds loses 150 requests during each 5-minute capped backoff gap.
- [17]
A startup probe budget shorter than real startup time turns a slow container into a killed container, which then re-enters the doubling CrashLoopBackOff schedule.
Sources
1 independent publisher whose own reporting we read for this story.
- How Kubernetes probes work
ngrok.com
1 article · August 22, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.