Skip to content

Build1 publisher2 min readPublished

Swarm services on Docker Engine 29.7 overlay networks start zero tasks on a host without IPv6

Docker Engine 29.7.0 and 29.8.2 started zero Swarm overlay tasks on a test host lacking IPv6, where 29.6.2 ran all 20, according to a dev.to post. It is one single-node reproduction, so Swarm teams leaving 29.6.x on similar hosts have reason to rerun it before upgrading.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Swarm services on Docker Engine 29.7 overlay networks start zero tasks on a host without IPv6
Generated illustration

What happened

  • On 29.8.2 every task failed at sandbox join with "overlay: cannot determine address family of transport: the local data-plane address is not currently known."
  • Rerunning 29.8.2 on the default VXLAN port 4789, with the older engine's overlay network torn down, produced the same error and the same 0/5 count.
  • A plain docker run container with a published port served nginx on the same 29.8.2 install, so the bridge driver and wider networking stack still worked.
  • The author did not bisect to a commit but points to a 29.7.0 and 29.8.0 release-note change that has service-mesh published ports share infrastructure with locally published ports.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Both post-29.6.x engines with reported results failed on this host, so for Swarm hosts without IPv6 the upgrade needs an overlay test service on a matching machine first.
  • cost Operators who want the 29.8.0 change that smooths overlay gossip CPU use would be taking an engine line that was still failing at 29.8.2 on this kind of host.
  • constraint The IPv6 warnings dockerd prints at startup also appear on the working 29.6.2, so logs cannot flag an at-risk host in advance; only a scheduled overlay task exposes the failure.

"I asked a Swarm service for twenty replicas on an overlay network and got zero," the author of a dev.to post wrote [1]. The author found it by accident. The author was setting up to measure a 29.8.0 change that spreads out the daemon's periodic overlay-network gossip so it does not burn CPU in bursts [2]. Every overlay service came up empty before any gossip was measured [2].

The test rig is careful work. Docker Engine 29.6.2 stayed as the system install [9]. Static builds of 29.7.0, 29.7.2 and 29.8.2 from download.docker.com each ran as a separate dockerd with its own data directory and Unix socket [9]. That let two versions run side by side on one host [9]. Each engine got a single-node docker swarm init, one overlay network and a service running sleep 600 [10]. Before trusting the 20/20 count on 29.6.2, the author pinged between two containers across the overlay and lost no packets [11].

Read literally, the error describes a missing input. To join a task to the overlay subnet, the driver needs to know whether its VXLAN transport is IPv4 or IPv6. The error says it went to the local data-plane address for that answer, and the address was not yet known [3]. The author concluded the fault is specific to the Swarm overlay driver's VXLAN data plane [15].

The IPv6 link comes from the host. It has no IPv6 sysctls and no /proc/net/if_inet6 table, and the author stressed that IPv6 was absent on it, not disabled [12]. That is a narrower case than a host with IPv6 turned off. Turning IPv6 off through the disable_ipv6 sysctl means writing to that file, and this host did not have the file [1]. Both 29.6.2 and every later engine log the same IPv6 read failures at startup, such as "failed to read ipv6 net.ipv6.conf.<bridge>.accept_ra" [13].

The post's headline says 29.7's overlay networking breaks every Swarm task without IPv6 [14]. The evidence behind it is one single-node host with no IPv6 stack [10][12]. I think the result is solid for hosts that match that one. For it to transfer to a production cluster, the cluster's hosts would need to lack an IPv6 stack and run Swarm tasks on an overlay network. Multi-node swarms and hosts with IPv6 switched off by sysctl were not tested [10][1].

What to watch

  • A moby issue or commit-level bisect that ties the failure to the change letting service-mesh published ports share infrastructure with locally published ports.
  • A 29.8.x or later release note that mentions the "cannot determine address family of transport" error.
  • A rerun on a multi-node swarm or a host with IPv6 present, to show whether the missing stack is the trigger.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories