Skip to content

Build1 publisher3 min readPublished

A double rollback in Docker's overlay driver cuts healthy containers off the network

Docker 29.8.2 cut two healthy containers off their overlay network within 20 seconds when a swarm service kept failing on a taken port, a dev.to test found. The containers keep showing as running, and after enough failures only a daemon restart brings them back.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Both containers lost carrier on their overlay interface and could reach nothing on the network, while docker ps kept showing them running.
  • In the swarm test, restarting or stop-starting the affected containers did not restore their traffic, and restarting the Docker daemon did.
  • The bug is filed upstream as moby/moby#53834, where it was reported against Docker 29.8.1.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure Any container sharing an overlay network with a crash-looping task on the same node can lose the network, and the only logged error belongs to the task that caused it.
  • constraint Process-only health checks keep passing through this failure, so detecting it requires a probe that sends traffic over the overlay or checks for the sandbox.
  • decision Once the count has gone well past zero, recovery means restarting the Docker daemon on the node, so a runbook that stops at restarting the container will not fix it.

The counter sits in the overlay driver, in `daemon/libnetwork/drivers/overlay/ov_network.go`. On 29.8.2 a join runs `n.joinCnt++`, and `leaveSandbox()` runs `n.joinCnt--`, then calls `n.destroySandbox()` if the count is zero [3]. The count is per network, per node, and zero is supposed to mean no endpoints are left [1]. The caller, `Endpoint.sbJoin` in `daemon/libnetwork/endpoint.go`, registers two rollbacks for one join [4]. A start that fails after the driver's Join adds one and removes two [2]. Every failure leaves the count one lower than it found it [1].

The write-up's test is small enough to rerun. It used a single-node swarm on a Debian 13 VM under Proxmox [6]. Two idle alpine containers sat on an attachable overlay network, and a third, on the default bridge, held host port 8080 [6]. A service then asked for 8080 in host mode. Each task failed with "Bind for 0.0.0.0:8080 failed: port is already allocated", and swarm scheduled another [7]. Ten seconds in, two tasks had failed and the sandbox was still there [8]. At 20 seconds, four had failed, the sandbox was gone, and pings from both containers to the network's load-balancer endpoint failed [8]. The retries ran at about one every five seconds [2].

How many failures it takes depends on how many endpoints are on the network. With the upstream report's trigger, an invalid endpoint sysctl, the author needed two failures to cut off one container and three for two [11]. Both runs needed one failure more than there were containers [3]. A retrying service supplies those failures with nobody touching the node [5]. Swarm is optional. With plain `docker run` and one container on the network, the second failed run against the taken port cut that container off [10].

The error and the damage show up in different places. The failing service logs a port conflict, and fixing the port lets it start [12]. It is the one thing on the node reporting a problem, and the only thing a port fix repairs [12][14]. The cut-off containers keep running. Their `eth0` shows `NO-CARRIER` and `state DOWN` [9]. Health checks that test only the process inside still pass, and `docker ps` and `docker inspect` still show them running [13]. Other services just see that they cannot reach c1 [13].

Recovery depends on how far past zero the count went, according to the write-up [15]. After only two manual failures on a one-container network, `docker restart c1` was enough [15]. After the swarm retries, restarting c1 and c2 left the sandbox missing [14]. Stopping both and starting them again recreated the namespace under `/run/docker/netns/`, but pings still failed [14]. Only `systemctl restart docker` brought both back [14].

The author's toolkit includes `check-overlay-sandbox.sh`, which finds containers in this state [16]. The signals from the test also work by hand: a missing namespace under `/run/docker/netns/`, `NO-CARRIER` on the overlay interface, and failed pings to the load-balancer endpoint [8][9]. I think any health check for an overlay-attached service has to send a packet across the overlay. A process check passes straight through this failure [13].

The write-up reproduced the bug on 29.8.2, the current release from Docker's apt repository [6]. The upstream report, moby/moby#53834, was filed against 29.8.1 [17]. The write-up does not say whether a fix is in progress. It is careful work. The reproduction is short, the cause is traced to two functions in the source, and the upstream trigger was rerun to confirm the same failure counts [3][4][11].

What to watch

  • A fix in moby/moby#53834 that makes the sbJoin rollback call Leave once, and which Docker release ships it.
  • Whether the same double rollback cuts off containers on multi-node swarms; the published test ran on a single node.
  • Reports of the bug on Docker releases before 29.8.1, which would widen the set of affected nodes.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories