Skip to content

Build1 publisher3 min readPublished

The database was fine: a one-in-three DNS failure hidden by an app that logs nothing

A WordPress pod failed to reach MySQL on a third of page loads. The resolver error only surfaced because the object cache in the same pod was noisier than the application.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • A WordPress site on the platform started throwing "Error establishing a database connection" intermittently, roughly one page load in three; reloading sometimes worked and sometimes did not.
  • The pod's config set WORDPRESS_DB_HOST = mysql.cdn.svc.cluster.local, so WordPress reached MySQL through a hostname rather than an IP.
  • MySQL was checked first: SHOW GLOBAL STATUS LIKE 'Connection_errors_max_connections' returned zero, there was no saturation and no slow queries, the server was healthy, and direct connections from outside the cluster worked.
  • WordPress does not log the database connection failure when WP_DEBUG is off, so the failure was invisible in the application log.
  • The same pod ran Redis for object caching, and Redis logged: "Redis::connect(): getaddrinfo for wp-redis-cache failed: Temporary failure in name resolution".

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A WordPress site on a Kubernetes platform began returning "Error establishing a database connection" on roughly one page load in three, and reloading fixed it [1]. According to a writeup on dev.to by the operators involved, MySQL was healthy the whole time: the real fault was a dead pod-network path to one of three CoreDNS replicas [12][15], and the reason it took so long to find is that the application swallowed the resolver error while another process in the same pod printed it in plain text [4][5].

The setup detail that decides everything: WORDPRESS_DB_HOST was set to mysql.cdn.svc.cluster.local, a hostname rather than an IP [2]. Every connection attempt therefore depended on a DNS lookup. The first wrong turn was the obvious one. SHOW GLOBAL STATUS LIKE 'Connection_errors_max_connections' returned zero, there was no saturation or slow queries, and direct connections from outside the cluster worked [3].

Nothing in the application log said otherwise, because WordPress does not log the database connection failure when WP_DEBUG is off [4]. That is the operational lesson worth keeping. The same pod also ran Redis for object caching, and Redis logged "Redis::connect(): getaddrinfo for wp-redis-cache failed: Temporary failure in name resolution" [5]. That is EAI_AGAIN, a DNS timeout [6], and since Redis also connects by hostname it was failing for the same reason as MySQL [7]. The author's rule: when a hostname-based dependency fails intermittently and silently, find a noisier hostname-based dependency in the same pod and read its log [19].

The second wrong turn was assuming CoreDNS itself was sick. All CoreDNS pods were Running and Ready, with no restarts and no errors in their logs [8]. There were three CoreDNS endpoints [9], and the failure rate was about one in three [1] - which is exactly what one dead endpoint out of three under round-robin produces [18]. Querying the Service ClusterIP cannot show you that, because it hides which backend answered and returns a mushy aggregate rate with no information about which one is bad [10].

The diagnostic that worked was mechanical: a busybox debug pod pinned to the affected node with nodeName, then 30 nslookups against each CoreDNS pod IP individually [11]. Two endpoints returned 30 of 30 successes; one returned 30 of 30 failures [12]. Because kube-dns round-robins across all three, a third of every lookup from that node went nowhere [13]. Repeating the same loop from a pod on a different node, the failing endpoint answered normally, which moved the fault from the pod to the link between two specific nodes [14].

The platform runs Weave, and Weave's control plane reported an established fastdp between those two nodes while pod-to-pod traffic across that link dropped 100 percent: a stale kernel datapath flow, reported healthy, forwarding nothing [15]. Deleting the weave-net pod on the affected node let the DaemonSet rebuild the datapath flows; it was back in about sixteen seconds, DNS went 30 for 30 from every endpoint, and the WordPress errors stopped [16].

Two things to watch. First, whether your own stack has a hostname dependency whose failures are invisible by default, since a silent resolver error turns a network fault into an application mystery [4]. Second, node-local DNS caching is the reflex hardening move, and in this case it was tried as a canary on one node and rolled back [17], so treat it as something to test per topology rather than adopt on faith. And treat overlay status output as a claim, not evidence: this link reported itself up while passing no traffic at all [15].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories