Build1 publisher3 min readPublished
The database was fine: a one-in-three DNS failure hidden by an app that logs nothing
A WordPress pod failed to reach MySQL on a third of page loads. The resolver error only surfaced because the object cache in the same pod was noisier than the application.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- A WordPress site on the platform started throwing "Error establishing a database connection" intermittently, roughly one page load in three; reloading sometimes worked and sometimes did not.
- The pod's config set WORDPRESS_DB_HOST = mysql.cdn.svc.cluster.local, so WordPress reached MySQL through a hostname rather than an IP.
- MySQL was checked first: SHOW GLOBAL STATUS LIKE 'Connection_errors_max_connections' returned zero, there was no saturation and no slow queries, the server was healthy, and direct connections from outside the cluster worked.
- WordPress does not log the database connection failure when WP_DEBUG is off, so the failure was invisible in the application log.
- The same pod ran Redis for object caching, and Redis logged: "Redis::connect(): getaddrinfo for wp-redis-cache failed: Temporary failure in name resolution".
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A WordPress site on a Kubernetes platform began returning "Error establishing a database connection" on roughly one page load in three, and reloading fixed it [1]. According to a writeup on dev.to by the operators involved, MySQL was healthy the whole time: the real fault was a dead pod-network path to one of three CoreDNS replicas [12][15], and the reason it took so long to find is that the application swallowed the resolver error while another process in the same pod printed it in plain text [4][5].
The setup detail that decides everything: WORDPRESS_DB_HOST was set to mysql.cdn.svc.cluster.local, a hostname rather than an IP [2]. Every connection attempt therefore depended on a DNS lookup. The first wrong turn was the obvious one. SHOW GLOBAL STATUS LIKE 'Connection_errors_max_connections' returned zero, there was no saturation or slow queries, and direct connections from outside the cluster worked [3].
Nothing in the application log said otherwise, because WordPress does not log the database connection failure when WP_DEBUG is off [4]. That is the operational lesson worth keeping. The same pod also ran Redis for object caching, and Redis logged "Redis::connect(): getaddrinfo for wp-redis-cache failed: Temporary failure in name resolution" [5]. That is EAI_AGAIN, a DNS timeout [6], and since Redis also connects by hostname it was failing for the same reason as MySQL [7]. The author's rule: when a hostname-based dependency fails intermittently and silently, find a noisier hostname-based dependency in the same pod and read its log [19].
The second wrong turn was assuming CoreDNS itself was sick. All CoreDNS pods were Running and Ready, with no restarts and no errors in their logs [8]. There were three CoreDNS endpoints [9], and the failure rate was about one in three [1] - which is exactly what one dead endpoint out of three under round-robin produces [18]. Querying the Service ClusterIP cannot show you that, because it hides which backend answered and returns a mushy aggregate rate with no information about which one is bad [10].
The diagnostic that worked was mechanical: a busybox debug pod pinned to the affected node with nodeName, then 30 nslookups against each CoreDNS pod IP individually [11]. Two endpoints returned 30 of 30 successes; one returned 30 of 30 failures [12]. Because kube-dns round-robins across all three, a third of every lookup from that node went nowhere [13]. Repeating the same loop from a pod on a different node, the failing endpoint answered normally, which moved the fault from the pod to the link between two specific nodes [14].
The platform runs Weave, and Weave's control plane reported an established fastdp between those two nodes while pod-to-pod traffic across that link dropped 100 percent: a stale kernel datapath flow, reported healthy, forwarding nothing [15]. Deleting the weave-net pod on the affected node let the DaemonSet rebuild the datapath flows; it was back in about sixteen seconds, DNS went 30 for 30 from every endpoint, and the WordPress errors stopped [16].
Two things to watch. First, whether your own stack has a hostname dependency whose failures are invisible by default, since a silent resolver error turns a network fault into an application mystery [4]. Second, node-local DNS caching is the reflex hardening move, and in this case it was tried as a canary on one node and rolled back [17], so treat it as something to test per topology rather than adopt on faith. And treat overlay status output as a claim, not evidence: this link reported itself up while passing no traffic at all [15].