Build1 publisher3 min readPublished
A hard-killed DigitalOcean Kubernetes node keeps receiving traffic for 10 seconds after it is marked NotReady
DigitalOcean Kubernetes marked a hard-killed node NotReady in 2.9 seconds and kept routing traffic to its pods for 10 more, across ten tests published on dev.to. On this managed platform the slow step is the endpoint update after detection, and none of the configuration changes tried shortened it.
The Engineer · Build desk

What happened
- Switching externalTrafficPolicy to Local looked like a four-fold improvement after three runs, but at five runs per policy the total losses were identical.
- In three of five Cluster-policy runs, 6 to 14 requests failed within a tenth of a second a minute or two later, while node readiness and the endpoint list stayed unchanged.
- After a planned kubectl drain, the endpoint list updated in 0.5 seconds and the run lost one request out of 31,652.
- The widely quoted figures are a 40-second node-monitor-grace-period default before NotReady and a five-minute default toleration before the taint manager evicts pods.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint Shortening node-controller timers would not cut loss on this platform. Detection already took under four seconds, and the dropped requests came after it.
- decision Picking a traffic policy means picking when requests fail: Local gives one bounded 10-second outage, while Cluster gives a smaller first hit and then later bursts.
- decision Planned maintenance should go through kubectl drain, the only path in these tests that pulled a node from the endpoints in about half a second.
- contradiction The Cluster bursts arrive a minute or two after the kill, outside the 10-second gap the author says holds every failure, so endpoint speed alone does not account for all the loss.
The author writes that every failure in the experiment happened inside the 10-second gap between NotReady and the endpoint update [7]. A drain was 26 times faster, which puts the unplanned endpoint update about 13 seconds after the power was cut [1]. Detection took 2.9 of those seconds [5]. At the test's rate of about 175 requests per second, roughly 1,750 requests arrive while the gap is open [2].
The instrumentation is what makes the numbers usable. A watcher sampled node readiness and the Service's endpoint list twice a second, so the cluster's view can be lined up against what the clients saw [4]. The first result surprised the author and the second contradicted it, so the experiment ran ten times [3].
The first fix tried was externalTrafficPolicy: Local [10]. Under the default Cluster policy, every node accepts traffic for the Service and forwards it to any pod. A surviving node keeps sending requests to pods on the dead machine until the endpoints catch up [10]. Under Local, a node serves only its own pods, and the load balancer's health check should drop the dead node within a few seconds [10]. That reasoning held up for three runs and fell apart at five [11]. "The setting did not change how much traffic I lost. It changed when I lost it," the author wrote [14].
Local took its whole loss in one window of 10.0, 10.0 and 10.1 seconds across three consecutive runs, then stopped [12]. The window is almost exactly as long as the endpoint gap [7]. The write-up does not say whether the health check or the endpoint update ended it. On the Cluster-policy bursts, the author argues that twenty threads failing together against a cluster whose state has not moved points into the load balancer [15]. I think those bursts are the one place in this data where the load balancer fails on its own.
The detection number comes from DigitalOcean's configuration. "I do not think the folklore is wrong so much as out of date," the author wrote, adding that the managed platform is clearly not running the stock timings [16][6]. A self-managed cluster left on the 40-second grace period should expect detection near the folklore figure [1]. The five-minute toleration could not show up in this data either. Each run observed only 150 seconds after the kill [3]. For the loss totals to transfer, a cluster needs detection tuned the same way and clients holding long-lived connections, like these 20 keep-alive clients on three 2-vCPU nodes [2]. Five runs per policy is a small sample, and the three-run conclusion has already reversed once [11].
The drain run's 31,652 requests equal a full three-minute run at 175 per second, which is about 31,500 [4]. One failure in a full run is about as cheap as taking a node out gets. Nothing the author changed in the configuration came close to the difference between planned and unplanned removal [9].
What to watch
- Whether the author pins down the load-balancer cause of the simultaneous Cluster-policy bursts.
- More runs per policy, since five runs already overturned the three-run conclusion about Local.
- A repeat on a self-managed cluster with the stock 40-second node-monitor-grace-period, to separate DigitalOcean's tuning from default Kubernetes behaviour.