Skip to content

Build1 publisher2 min readPublished

Route 53 health checks that need TLS 1.2 pulled traffic from a TLS 1.3-only region

Disabling TLS 1.2 on public ingress failed Route 53 HTTPS health checks and cut traffic to a healthy region, according to an InfoQ case study. The fault sat in the control plane that decides where traffic goes, where no in-region dashboard could see it.

The Engineer · Build desk

Illustration accompanying Route 53 health checks that need TLS 1.2 pulled traffic from a TLS 1.3-only region

What happened

  • The team moved its public ingress load balancers to TLS 1.3 for compliance, and handshakes, service health and metrics all looked normal afterwards.
  • Users were sent to a region thousands of kilometres away, where latency spiked and the failover region scaled under load it had not been provisioned for.
  • Synthetic monitoring eventually exposed a pattern internal telemetry could not show, and isolating the cause took roughly forty minutes.
  • The author argues that multi-region designs still share control-plane dependencies, including DNS health checks, identity and access management, and routing.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Any TLS policy change on public ingress has to be tested against the provider's health checker as a client, because user clients and that probe negotiate TLS separately.
  • exposure Alerting built on in-region service metrics stays quiet through a routing failure, so detection depends on probes that check whether traffic arrives at all.
  • decision Each failover path needs a named owner and a test schedule; the author argues that on-call ownership alone does not keep recovery procedures current.
  • cost Full failover drills are expensive and risky, so organisations that skip them are relying on recovery paths they have never seen work.

Route 53's health checker is a TLS client, and it has its own requirements. According to the article, Route 53 HTTPS health checks required the target endpoint to support TLS 1.2 at the time of writing [2]. When an endpoint offers only TLS 1.3, the checker cannot complete the handshake and marks the endpoint unhealthy [2]. Compliance checklists seldom list the cloud provider's own health checker as a client. The handshakes the team watched after the rollout all completed, and nothing in the rollout showed the probe failing [1].

The failure was hard to detect because of where it happened. It sat in the control plane, the layer that decides where traffic goes, while the application data plane kept working [7]. Service metrics describe the data plane. A region can pass every data-plane check while receiving no traffic at all [4].

The author separates two terms that teams tend to use as synonyms. High availability means surviving expected failures with minimal interruption. Resilience means recovering from conditions the system was never designed to handle [14]. Only the first comes with easy metrics: uptime, failover timing, replication lag [14]. Multi-AZ designs model a whole zone going down, the article notes. They do not model correlation at the software layer, such as a bad config pushed to every replica at once, a poisoned cache record served from every read replica, or a dependency upgrade that silently breaks a contract [12]. The TLS change was the third kind. The contract it broke was with a health-check service that every region depended on [8].

"A failover path that has never been exercised is not a recovery strategy; it is just an assumption," the author wrote [11]. The article calls resilience probabilistic. Its stated goal is confidence, built by testing recovery paths again and again until hidden dependencies show up [13]. I think this incident also makes the case for a test much smaller than a full regional failover. Point a health check of the same type at a staging listener running the new TLS policy. That would have exercised the exact dependency that failed, and no production traffic would have moved.

The account comes from one practitioner. The article does not name the company, the date of the incident or the CDN involved. The TLS 1.2 requirement is also stated "as of this writing" [2]. It describes Route 53 at one point in time, and anyone designing around it would need to check it against current Route 53 documentation.

What to watch

  • An update to Route 53 HTTPS health checks that lets them complete a TLS 1.3-only handshake would close this specific failure path.
  • More incident write-ups that name IAM or routing as the shared dependency would show whether the article's control-plane argument holds beyond a single TLS case.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories