Skip to content

Build1 publisher3 min readPublished

Envoy Gateway's 1.9.1 timeout revert needs a fast proxy rollout on clusters running 1.9.0

Envoy Gateway 1.9.1 puts the SDS and RDS initial fetch timeout back to 15 seconds after 1.9.0 set it to zero. Clusters on 1.9.0 have to plan the upgrade around a fast proxy replacement that needs spare capacity and can drop long-lived connections.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Envoy Gateway's 1.9.1 timeout revert needs a fast proxy rollout on clusters running 1.9.0
Generated illustration

What happened

  • A production user reported on GitHub an outage in which long-running proxies lost their downstream TLS certificates and stopped serving HTTPS until restarted by hand.
  • The project tells clusters on 1.8.x to upgrade straight to 1.9.1 and skip 1.9.0 altogether.
  • The release also moves OAuth2 and OIDC session cookies to AES-256-GCM and drops AES-256-CBC decryption, addressing CVE-2026-47775.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision The starting version now sets the upgrade plan: 1.8.x clusters keep the timeout behaviour they already had, while 1.9.0 clusters must plan a proxy rollout alongside the controller upgrade.
  • cost On 1.9.0 clusters, one maintenance window can both interrupt long-lived WebSocket and gRPC streams and force users on old-format sessions to log in again.
  • exposure Post-upgrade checks that test only front-door HTTPS miss two places the same certificate or CA failures can appear: backend TLS and the global rate-limit service.
  • constraint Teams running a replacement Envoy bootstrap carry an extra manual step, since the cookie-encryption runtime settings have to be configured explicitly.

In Envoy Gateway 1.9.0 the initial fetch timeout for SDS and RDS was zero [2]. According to the project, that meant a cluster waiting on a Secret or endpoint that never arrived waited indefinitely and stayed in a warming state [3]. Zero, in this setting, is the longest timeout on offer. While the cluster warmed, CDS updates stopped progressing and health checks were delayed [3]. Version 1.9.1, released on August 28 [1], goes back to Envoy's default of 15 seconds, the model 1.8.x used [4].

For a 1.8.x cluster that goes straight to 1.9.1, the timeout behaviour does not change at all [2]. A cluster already on 1.9.0 is different. There, the controller upgrade changes the SDS configuration under proxies that are still serving traffic. Envoy Gateway warns that this can trigger an Envoy issue on proxies that keep running while the new controller deploys [5]. TLS listeners on those proxies can go active without their certificates, and new TLS handshakes fail [6]. Backend TLS configurations and the global rate-limit service can hit similar certificate or CA problems [7].

The failure has already shown up in production. A GitHub issue described by InfoQ reports an outage in which long-running Envoy proxies lost their downstream TLS certificates and stopped serving HTTPS, and recovery took manual proxy restarts [11]. The issue ties the behaviour to how new SDS subscriptions are handled and to the timeout change in 1.9.0 [12].

For 1.9.0 clusters, the project recommends a rolling-update configuration that replaces proxy pods quickly [9]. The warning is about proxies that stay up while the new controller rolls out [5]. I think the fast rollout exists to cut how long any old proxy runs against the new SDS configuration. It has a price. It needs enough cluster capacity for the replacement pods, and existing connections can be terminated as pods are replaced [9]. The release notes single out long-lived WebSocket and gRPC connections as the ones that can be interrupted [10].

The order I would use on a 1.9.0 cluster:

1. Confirm there is spare capacity for a fast proxy replacement [9]. 2. Set the rolling-update configuration before the controller upgrade starts, so proxy replacement follows the controller closely [9]. 3. Upgrade the controller. 4. Test new TLS handshakes on downstream listeners, on backend TLS and on the rate-limit service [6][7].

The same upgrade breaks some sessions. 1.9.1 encrypts OAuth2 and OIDC session cookies with AES-256-GCM and removes the legacy AES-256-CBC decryption path, closing the padding-oracle flaw tracked as CVE-2026-47775 [13]. Sessions encrypted the old way have to authenticate again after the upgrade [14]. On a 1.9.0 cluster, dropped streams and forced logins land in the same window [1]. Deployments that use a replacement Envoy bootstrap have to configure the matching runtime settings themselves [15].

The rest of the security work is careful. Before 1.9.1, an OCI registry that rejected an HTTPS request could push Envoy Gateway onto plain HTTP, and an on-path attacker could then supply arbitrary Wasm code [16]. HTTP is now allowed only for registries explicitly configured as insecure [17]. A tenant-supplied Kubernetes security context in EnvoyProxy configuration can also no longer fully replace Envoy Gateway's hardened defaults, a gap that could have removed restrictions on privileged execution or running as root [18].

What to watch

  • Whether upstream Envoy fixes the SDS subscription issue so running proxies can take the configuration change without being replaced.
  • Further reports on the GitHub outage issue from 1.9.0 operators who upgrade without a fast proxy rollout.
  • Whether a later 1.9.x patch changes the recommended path for clusters already on 1.9.0.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories