Build1 publisher3 min readPublished
Deleting kube-proxy moves your service traffic where your SIEM cannot follow it
One operator flipped Cilium to full kube-proxy replacement and woke at 2:47 AM to a SIEM ingesting zero events. The cluster was healthy. The evidence pipeline was not.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- The author ran a Cilium upgrade on a four-node bare-metal homelab cluster consisting of a Dell OptiPlex and three Raspberry Pis, running Talos Linux, Cilium, ArgoCD and Longhorn.
- Before the change the cluster ran Cilium 1.15 in kube-proxy-replacement partial mode, with kube-proxy's iptables rules still doing the heavy lifting for NodePort and ClusterIP services.
- The Cilium 1.16 release notes described the eBPF kube-proxy replacement as "production-ready for all environments".
- The Talos Linux docs covered the migration with a single bullet point: "Set kubeProxyReplacement: true in Cilium values."
- The change was made with a helm upgrade of Cilium in namespace kube-system setting kubeProxyReplacement=true, k8sServiceHost=auto and k8sServicePort=6443.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
An operator upgraded Cilium on a four-node bare-metal cluster from partial kube-proxy replacement to full replacement, and the alert that eventually fired came not from the cluster but from his SIEM, whose ingestion rate for the `kubernetes-audit` and `network-flow` streams had flatlined to zero at 2:47 AM [1][2][7]. Every health signal stayed green through it, which is why this belongs in the change-review conversation and not the bug tracker [6][17].
The setup is a homelab: a Dell OptiPlex plus three Raspberry Pis running Talos Linux, Cilium, ArgoCD and Longhorn [1]. Before the change, Cilium 1.15 was in partial mode, with kube-proxy's iptables rules still doing the work for NodePort and ClusterIP [2]. The Cilium 1.16 notes called the eBPF replacement "production-ready for all environments," and the Talos documentation reduced the migration to one bullet point: set `kubeProxyReplacement: true` [3][4]. The actual change was a three-flag `helm upgrade` [5]. Pods rolled, nodes stayed up, `kubectl get pods -A` looked normal, and the author went to bed [6].
The mechanics explain the blindness. kube-proxy watches the API for Service and EndpointSlice changes and writes them into iptables as PREROUTING chains, a KUBE-SERVICES target and a DNAT to a backend pod IP [12]. Cilium instead attaches eBPF programs at `tc` hooks, looks the ClusterIP up in a BPF map keyed on address and port, rewrites the destination and forwards [14]. The author's own summary of the new path is "no iptables, no chains, no table reloads" [15]. That is the performance case and the observability liability in the same sentence: if your flow logging, host IDS or network agent reads conntrack tables or netfilter counters, the surface it reads has been removed. He calls the post an autopsy of deleting the datapath a security stack depends on, and by his headline the gap ran six hours [18][8].
What makes this worth attention is that observability was one of the stated reasons for the migration. iptables drops packets silently, with no log line and no metric, and identifying the offending rule means running `iptables -L -v -n` on every node [11]. Hubble was supposed to fix that with per-flow visibility [16]. It did, in the sense that Hubble showed flows after the cut-over [17]. But flows in Hubble are not events in the SIEM, and nothing in the one-bullet migration guide connects the two.
The performance case was real, at least at this scale: roughly 120 services generated about 3,500 iptables rules across the nodes, or around 29 rules per service, with every new service triggering a full `iptables-restore` that locked netfilter for hundreds of milliseconds [9][19][10]. On a Raspberry Pi that is a visible latency spike [10]. The old path also carried O(n) per-packet rule evaluation and no connection tracking visibility without conntrack tools [13].
Two things to check before you make this change. First, enumerate which of your log and metric streams actually derive from netfilter or conntrack, and confirm each has a Hubble-based or agent-based equivalent that is already exporting off-cluster. Second, treat SIEM ingestion rate as a cluster health signal, because in this account it was the only thing that noticed [7][17]. One caveat on the source: the published excerpt breaks off before naming the root cause or the fix, so the diagnosis is not available to copy [20].