BuildNot yet confirmed elsewhere1 publisher2 min readPublished
Kubernetes spreads pods where it likes; a few manifest lines and some arithmetic decide the rest
Gremlin's walkthrough of topology spread constraints is sound mechanics. Its own worked example still permits three of four replicas in one region, and the setting it recommends cannot enforce anything.
The Engineer · Build desk
What happened
- Gremlin published a walkthrough arguing that Kubernetes placement is not reliability-aware by default and that a few manifest lines make a deployment zone-redundant.
- Topology spread constraints sit in spec.topologySpreadConstraints, group nodes into domains by a node label, and can be set per workload or as a cluster-wide default.
- whenUnsatisfiable defaults to DoNotSchedule; ScheduleAnyway demotes the rule to a scheduling preference that merely favours less skewed topologies.
- In the unconstrained example, Gremlin says the scheduler may put three of four replicas in us-east-1a, or all four on one node.
- The recommended manifest uses maxSkew 1, the zone topology key, ScheduleAnyway, and a labelSelector matching app: nginx.
Why it matters
- constraint Four replicas over three zones under maxSkew 1 has exactly one legal shape, 2/1/1, so even the tightest spread concedes half the fleet to a single zone loss.
- contradiction Gremlin recommends ScheduleAnyway while describing maxSkew 1 as ensuring balance.
- decision Replica count stops being purely a capacity knob and becomes a blast-radius setting: pick a multiple of your domain count or accept a fat domain that decides how much you lose.
- exposure Anyone spreading on zone across a cluster whose zones cluster into one region is exposed at the region level while passing every zone check, since two of the three example domains share us-east-1.
maxSkew bounds a difference. It does not distribute anything. With three zones and four replicas, the tightest legal setting still has to put the fourth replica somewhere, and 2/1/1 is the only shape available [7][1][13]. Lose the zone holding two and half the replicas go with it [13]. That is the arithmetic inherited by every capacity plan resting on the phrase "we are zone-redundant".
The region detail in Gremlin's own example is the sharper one. Two of the three zones, us-east-1a and us-east-1b, sit in the same region [7]. Spreading on `topology.kubernetes.io/zone` treats them as two independent domains, so a perfectly legal 2/1/1 with the pair in us-east-1a puts three of four replicas inside us-east-1 [9]. That is the concentration the post opens by warning against [8]. A zone key buys zone redundancy and nothing above it.
Then there is `whenUnsatisfiable`. Gremlin sets it to `ScheduleAnyway` so the pod runs even when the constraint cannot be met [1], and describes `maxSkew: 1` in the same passage as ensuring that no node carries more than one extra replica [15]. Both cannot be true. `ScheduleAnyway` is a scoring preference that favours topologies which reduce skew [5]; under real zone pressure, from a drain or a capacity shortfall, the scheduler packs pods into the surviving domains and reports success. `DoNotSchedule`, the default [5], is the version that means what it says, and its price is Pending replicas at exactly that moment. Which failure you prefer is the actual decision, and it is buried in one word of YAML.
Worth noticing too that the recommendation counts nodes in prose while keying on zones in the manifest [1][15]. A slip in a blog post costs nothing. The same slip written with `topologyKey: kubernetes.io/hostname` gives you tidy node-level balance and no zone guarantee at all.
One structural detail in the field list repays attention: the restriction is one constraint per `topologyKey` and `whenUnsatisfiable` pair [6], which leaves both values of `whenUnsatisfiable` available on the same key [14]. A hard rule at a loose skew alongside a soft rule at a tight skew is therefore expressible. Most manifests carry one constraint and stop there.
Which leaves the question the material does not answer. The text we have ends at a heading promising how to find pods with missing topology spread constraints [16], and at fleet scale that is the whole problem: a cluster where four deployments set the field and forty omit it has no spread policy, it has four opinions. Cluster-level defaults exist [3], and they are the only version of this that survives contact with a platform team that does not review every manifest.
What to watch
- Whether Gremlin publishes the promised method for finding workloads with no spread constraints, the step that turns this from a manifest tip into a fleet inventory.
- How a cluster-level default constraint resolves against a workload declaring its own rule on the same topologyKey, which the post leaves open and which decides who owns spread policy.
- Whether teams running DoNotSchedule report Pending replicas during genuine zone capacity crunches, the cost that configuration quietly accepts.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence54
- Adoption
- Insufficient
- Hype gap+38
- Incentives72
- Confidence56
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The recommended manifest sets maxSkew to 1, topologyKey to topology.kubernetes.io/zone so spread is limited by zone even where zones span multiple regions, whenUnsatisfiable to ScheduleAnyway so the pod runs even if the constraint cannot be satisfied, and a labelSelector matching app: nginx.
- [2]
Gremlin's post argues that most Kubernetes users pay little attention to where pods are placed, that pod distribution plays a bigger role in reliability than they think, and that adding a few lines to a manifest can make deployments zone-redundant and evenly scalable.
- [3]
Topology spread constraints determine how Kubernetes distributes pods across failure domains such as regions, zones and nodes. They are defined in the field spec.topologySpreadConstraints and can be applied to a pod or set at cluster level as a default.
- [4]
The constraint fields are maxSkew, minDomains, topologyKey, whenUnsatisfiable, labelSelector, matchLabelKeys, nodeAffinityPolicy and nodeTaintsPolicy. topologyKey is the node label whose values group nodes into topology domains; minDomains is the minimum number of eligible domains.
- [5]
whenUnsatisfiable defaults to DoNotSchedule, under which maxSkew is the maximum permitted difference between the minimum number of pods in a domain and the matching pods in the target topology. Setting ScheduleAnyway schedules the pod regardless, giving higher precedence to topologies that reduce the skew.
- [6]
Only one topologySpreadConstraint may be defined for a given topologyKey and whenUnsatisfiable pair.
- [7]
The worked example uses a cluster spanning three availability zones, us-east-1a, us-east-1b and us-west-2a, with four replicas of one pod deployed for redundancy.
- [8]
Without constraints, Gremlin says the scheduler might place two pods on two nodes and leave one empty, or three pods in us-east-1a and one in us-east-1b, which puts the deployment at risk if the us-east region goes down, or in the worst case place all four on one node and create a single point of failure.
- [9]
Because two of the three named zones (us-east-1a, us-east-1b) are in the us-east-1 region, a legal 2/1/1 spread with the two-pod domain in us-east-1 places three of four replicas, 75 percent, in that single region.
- [10]
Under maxSkew 1 the fullest domain holds ceil(replicas / domains), so the spread is even only when the replica count is a multiple of the domain count: six replicas over three zones gives 2/2/2 and a 33 percent loss per zone, while four gives 50 percent.
- [11]
nodeAffinityPolicy defaults to Honor, which limits the topology calculation to nodes matching the pod's nodeAffinity and nodeSelector; Ignore uses all nodes. nodeTaintsPolicy decides whether node taints are included in the calculation.
- [12]
Gremlin's round-robin illustration of four replicas produces one node with two pods and two nodes with one pod each.
- [13]
Four replicas across three zone domains under maxSkew 1 can only be arranged 2/1/1, so the loss of the fullest zone removes two of four replicas, or 50 percent.
- [14]
Since the limit is one constraint per topologyKey and whenUnsatisfiable pair, two constraints on the same topologyKey are permitted provided they use different whenUnsatisfiable values.
- [15]
The post states that setting maxSkew to 1 ensures that no single node has more than one additional replica of the pod than any other node, in the same walkthrough that sets topologyKey to zone and whenUnsatisfiable to ScheduleAnyway.
- [16]
The supplied text of the post ends at a heading, "How to find pods with missing topology spread constraints", without the method itself.
Sources
1 independent publisher whose own reporting we read for this story.
- Optimizing Kubernetes pods for reliability with topology spread constraints
gremlin.com
1 article · August 23, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.