Build1 publisher3 min readPublished
Kubernetes' default rollout settings can stall single-replica apps on ReadWriteOnce volumes
Professional IT Services reviewed all 15 volume-backed Deployments on its K3s cluster after a routine node upgrade caused a three-hour outage. Three now use Recreate, because the default rolling update can stall on a ReadWriteOnce volume.
The Engineer · Build desk

What happened
- If that new pod lands on another node, the volume is still attached to the old one, and the pod waits in ContainerCreating with a Multi-Attach error.
- Every volume-mounting Deployment on the cluster runs one replica on a ReadWriteOnce Hetzner volume, including a mail stack whose five components share 100 GiB.
- The WordPress sites kept zero-downtime rolling updates through a preferred pod-affinity rule that draws each new pod to its predecessor's node.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Every single-replica app holding a ReadWriteOnce claim needs its strategy picked deliberately: a few seconds of downtime per rollout under Recreate, or same-node placement under RollingUpdate.
- exposure The WordPress fix rests on a scheduler preference, so any replacement placed on another node falls back into the Multi-Attach stall the review set out to remove.
- constraint Moving those volumes to ReadWriteOncePod would close off the same-node rollout, because that mode admits one pod where ReadWriteOnce admits one node.
Professional IT Services wrote that the standard three-line guide, Deployments for stateless apps, DaemonSets for one pod per node and StatefulSets for databases [1], "is correct, and it is also the part that never causes an outage" [2]. In the post's account, outages come from how each controller behaves when a node is drained, a volume refuses to attach, or a rollout meets a pod that never turns healthy [3].
Its cluster is small enough to reason about. Five K3s nodes on Hetzner Cloud run 50 Deployments, 6 DaemonSets and 12 StatefulSets, with 26 persistent volumes [4]. Fifteen of those Deployments, 30%, mount a claim [9][1].
RollingUpdate's defaults suit interchangeable pods. Surge and unavailability are both 25%, with surge rounded up and unavailability rounded down [6]. A quarter of one replica is 0.25, so the surge becomes 1 and the unavailability 0 [3]. The new pod starts before the old one stops [7]. If the scheduler puts it on another node, it sits in ContainerCreating with a Multi-Attach error, waiting for a volume still attached to the old node, and the rollout stalls [8].
The fix depends on what ReadWriteOnce actually restricts. It limits a volume to one node at a time; the per-pod mode is ReadWriteOncePod [12]. Two pods on the same node can both mount a ReadWriteOnce volume [12]. The post also corrects its own earlier version, which had said Deployments have ephemeral storage. A Deployment can mount a PersistentVolumeClaim like any other pod [15].
After the outage, the team settled on a rule, in the post's words: "Recreate where it is necessary, RollingUpdate where it is feasible" [10]. Recreate stops the old pod before starting the new one. Each rollout costs a few seconds of downtime, and it cannot deadlock on a volume [11]. Three of the 15 use it, leaving 12 that do not [11][2]. The WordPress sites kept zero-downtime rollouts by staying on one node [13]. Their chart does it with a `pod-group` label and a single affinity term: `preferredDuringSchedulingIgnoredDuringExecution`, weight 100, `topologyKey: kubernetes.io/hostname` [14].
I like this design. One scheduling term uses the access mode exactly as specified. It is also a preference, as the key says, and the post describes the new pod as drawn to its predecessor's node [14]. A replacement the scheduler places elsewhere anyway meets the same Multi-Attach stall [8].
Reproducing the failure on another cluster takes three conditions: one replica, a ReadWriteOnce claim, and a scheduler free to pick a different node [7][8]. A single-node cluster has nowhere else to put the pod.
The post says the workload controllers decided both how bad the 17 June outage got and how it was repaired [5]. The Deployment half of that claim is laid out in detail. The drain sequence, the repair, and what the six DaemonSets and 12 StatefulSets did during the upgrade are not described in the available text [4].
What to watch
- The rest of the post's account of 17 June: how the six DaemonSets and 12 StatefulSets behaved during the drain would test the claim that controller choice set the outage's severity.
- Whether the 12 volume-backed Deployments not on Recreate keep same-node rollouts when the preferred node is the one being upgraded.
- Whether the team hardens the WordPress affinity from preferred to required scheduling.