Build1 publisher3 min readPublished
Static pods led a hand-built Kubernetes controller to report the whole control plane at risk
One Karpenter operator wrote about 250 lines of plain client-go to flag pods that will not come back when a tainted node goes away. Doing it without controller-runtime or Kubebuilder put every piece of the watch, queue and reconcile loop on the page.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The controller sorts pods by owner: ReplicaSet, StatefulSet and Job pods reschedule, DaemonSet pods go with their node, and pods with no controller owner are lost for good.
- Its first run against a node tainted spot-interruption=true:NoSchedule reported etcd, kube-apiserver, kube-scheduler and kube-controller-manager as bare pods that would not be recreated.
- The flagged pods were static pods, which the kubelet restarts from manifest files on disk whatever the API records about their owners.
- A second version wrote an at-risk annotation onto each flagged pod, and the annotation stayed on after the taint was removed and the node was healthy again.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- exposure An operator who answers kubectl drain's refusal by adding --force deletes the ownerless pods for good, so a warning only helps if it arrives at taint time, before anyone types that flag.
- constraint An at-risk check that relies only on GetControllerOf returning nil lists every static control-plane pod, so it needs the mirror-pod filter before anyone can act on its output.
- decision A controller that writes findings onto pods has to do work when a node is healthy as well; one that skips that branch leaves annotations that drift away from what is true on the node.
"Everything left of reconcile is plumbing," the author wrote [7]. The diagram in the post lays that plumbing out [6]. One WATCH is opened against the API server and held open, and a node informer keeps a local cache [6]. The informer's add, update and delete handlers push the node name onto a workqueue [6]. The workqueue dedupes keys, rate-limits them and tracks what is in flight [6]. A worker goroutine pulls each key and calls reconcile [6]. The decisions fit in about 40 of the roughly 250 lines, around 16 percent of the program [7][1].
The author set out to explain what Karpenter does between "pod is unschedulable" and "EC2 instance appears" [1]. This controller does not answer that question [3]. It is a smaller controller that watches nodes for a disruption taint, and the author's case for building it is that its shape "is the same as every controller" [3][8]. The part that carries over to Karpenter is that loop [8]. The post does not examine Karpenter's own code [8].
The first bug is the one I would want anyone writing drain tooling to read. The check behind the four control-plane warnings was `metav1.GetControllerOf(pod)` returning nil, and nil was correct: those pods have no controller owner in the API [9][10]. A tool that puts etcd on its casualty list gets read carefully [9]. For a static pod, the API holds a read-only mirror that the kubelet creates so the pod appears in `kubectl get pods`, with the Node set as its owner [11]. The kubelet works from the manifest on disk, so the ownership record in the API does not decide whether the pod comes back [11]. The author's lesson was that "a correct check can produce a wrong answer" [13]. The fix is a `continue` for any pod carrying the `kubernetes.io/config.mirror` annotation, which the kubelet sets on every mirror pod [12].
The line behind the stale annotation was `if taint == nil { return nil }` [15]. "I'd written half a reconcile loop," the author wrote [16]. The repaired branch returns `c.clearAnnotations(name)` [15]. The author's definition of reconcile is "make reality match what it should be, whatever reality currently is" [17]. A node with no taint is healthy, and the loop has to enforce that state as well [20]. The cleanup keys off the annotation the controller left behind, not the conditions it used when writing it, and the author ties that choice to later changes in which pods get flagged [18]. I think keying on the marker is the most careful decision in the post, because a pod marked under an old rule still gets cleared under a new one [18].
The controller reports. It does not block a drain or evict anything itself [3]. The author wanted ownerless pods surfaced before the drain, "which is when it's actually useful" [19]. Whether that prevents a permanent loss depends on someone reading the annotation before the drain runs [3][5].
What to watch
- Publication of the controller's source, so others can test the 40-line reconcile against static pods and recovered nodes.
- Any move from annotating pods to delaying or blocking a drain, so that the report stops anything instead of only warning.
- Whether the classification gains a capacity check, since the post counts any ReplicaSet, StatefulSet or Job pod as fine once its node goes.