Skip to content

Build1 publisher3 min readPublished

Cross-node pod traffic: routing or encapsulation, and why that choice is a debugging decision

A dev.to walkthrough traces one packet from 10.244.1.5 to 10.244.2.5. The operationally useful part is not the hop count, it is which route table or outer header you get to read.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Cross-node pod traffic: routing or encapsulation, and why that choice is a debugging decision
Generated illustration

What happened

  • The article is Level 2 of a Kubernetes networking series on dev.to; Level 0 covered core Linux networking primitives and Level 1 covered how a pod gets its own network namespace, eth0 interface and IP address. Level 2 promises a full hop-by-hop trace of cross-node pod traffic including routing, overlay networking, VXLAN and BGP.
  • The worked scenario is Pod A with IP 10.244.1.5 on node-1 and Pod B with IP 10.244.2.5 on node-2, with Pod A sending traffic 10.244.1.5 to 10.244.2.5.
  • The article states the Kubernetes networking model as: every pod should be able to communicate with every other pod, without anyone needing to manually manage routes for each pod.
  • Every pod has its own routing table; the example given is 'default via 10.244.1.1 dev eth0' and '10.244.1.0/24 dev eth0'. When Pod A targets 10.244.1.6, Linux finds it falls within 10.244.1.0/24 and sends the packet out eth0.
  • For two pods on the same node the conceptual flow is Pod A, eth0, veth, node networking, veth, eth0, Pod B; the article says the exact implementation depends on the CNI plugin but the core idea (Pod A to node networking to Pod B) holds.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A Kubernetes networking series on dev.to published its Level 2 installment, which traces cross-node pod traffic hop by hop using a concrete pair: Pod A at 10.244.1.5 on node-1, Pod B at 10.244.2.5 on node-2 [1][2]. The reason this matters past tutorial value is that the same trace splits into two very different failure surfaces depending on whether your CNI teaches the underlying network about pod subnets or wraps pod packets inside node-to-node packets [10].

Start with the easy case, because it sets the baseline. On a single node the pod's own routing table does the work: a default via 10.244.1.1 out eth0, plus a directly connected 10.244.1.0/24, so a packet to 10.244.1.6 matches the local subnet and leaves through eth0 [4]. From there it is veth to node networking to veth to the peer's eth0 [5]. The article notes the exact mechanics vary by CNI plugin, but that shape holds [5].

Cross the node boundary and the pod's table stops being interesting. The packet, source 10.244.1.5 and destination 10.244.2.5, exits eth0, crosses the veth pair, and lands in node-1's stack, which now has to answer where 10.244.2.5 lives [6]. The conceptual node table is one line per peer: 10.244.1.0/24 local, 10.244.2.0/24 via Node 2 [7]. That is the general form of every cross-node pod flow, Pod A to Node 1 to a route to Node 2 to Pod B [8]. The consequence for on-call work is that the deciding piece of state is not visible from inside the pod at all [3].

Who fills that table in is the CNI's problem, and the article's stated job description for a CNI is simply to make sure the network knows how to reach every pod IP [9]. It lists the mechanisms in play: routing, overlay networking, VXLAN, Geneve, BGP, cloud-native routing, and eBPF [9]. Those collapse into two strategies. Either the physical or cloud network is taught the pod networks directly, giving Pod A to Node 1 to a router to Node 2 to Pod B, or the pod packet is encapsulated so the outer header reads Node1-IP to Node2-IP while carrying 10.244.1.5 to 10.244.2.5 inside [10][11]. The article puts BGP directly after its section on routing without an overlay, that is, as the way those pod-subnet routes get distributed [16].

The debugging split follows mechanically. On the encapsulated path, a capture taken on the node uplink shows node addresses in the outer header, and the pod addresses exist only in the payload, so anything you filter or ACL on the underlay is filtering node traffic [1]. On the plain-routed path, the pod address is what the underlay forwards on, so the artifact you read is the route table and whatever advertised it, not a decapsulated inner frame [2].

Two things worth pinning down for your own clusters. First, the article is explicit that Kubernetes does not mandate an overlay [13], so this is a per-cluster fact you look up rather than assume, and the model only promises that every pod can reach every pod without anyone hand-managing per-pod routes [3]. Second, pod IPs are not internet-routable [14], which is where the series goes next, into Services [15].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories