Skip to content

BuildNot yet confirmed elsewhere1 publisher2 min readPublished

Identical Helm charts, three clouds, one OOMKill loop: portability is a claim about YAML

The chart was identical on EKS, AKS and GKE. Only GKE's nodes had cgroup v2, and the JVM inside sized its heap from the host's memory instead of the container's.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Identical Helm charts, three clouds, one OOMKill loop: portability is a claim about YAML
Generated illustration

What happened

  • One Helm chart, identical values and the same image tag, went out to EKS, AKS and GKE inside a single 20-minute release window.
  • EKS and AKS came up healthy; the GKE pods ran for about 90 seconds and then OOMKilled in a loop.
  • GKE's newer node images had moved to containerd with cgroup v2, while the EKS and AKS nodes were still on cgroup v1.
  • The service's older JVM read 64GB of host memory instead of the 2GB container limit and allocated an 8GB heap.
  • Four hours went into finding the cause.

Why it matters

  • constraint A memory limit is enforceable across all three providers, but the interface a process reads to discover that limit is not, and it changes when a provider rolls a node image.
  • decision Anything that self-tunes from the machine it lands on now needs explicit sizing or pinned node images per pool, because leaving the runtime to guess makes a vendor release train a production...
  • cost Two of the three clusters could not reproduce the fault, so pre-production on either would have passed clean and the bill was paid in incident hours instead.
  • exposure The same silent-fallback shape shows up in identity: a mistyped IRSA annotation yields a working pod running on the node's permissions, with nothing in the chart or the logs to show it.

The failure is a path lookup. Under cgroup v1 a container's memory ceiling sits at `/sys/fs/cgroup/memory/memory.limit_in_bytes`; under v2 the same number lives at `/sys/fs/cgroup/memory.max` [6][7]. A JVM that only knows the first path does not fail loudly when the file is absent. It falls back to what the host advertises, which here was 32 times the ceiling the container actually had, and then sized a heap four times that ceiling [15][16].

Worth separating what travelled from what did not. The limit travelled: GKE's kernel enforced exactly what the chart asked for, which is precisely why the process was killed instead of quietly growing past it [8]. What no chart can express is which file the process will read to learn that figure. That is a property of the node image, and node images ship on the provider's calendar rather than yours [5].

Put the clock against the loop. At roughly 90 seconds per crash, four hours of diagnosis works out to something like 160 restart cycles, every one of them identical and none of them informative [18].

The networking chapter is the same defect in different clothes. EKS with the AWS VPC CNI hands every pod a routable address from the VPC subnet [10], and a t3.medium tops out at 3 ENIs times 6 IPs, so 18 addresses, 17 of them available to pods [11][17]. Ask a /24 for 200 replicas and pods sit Pending on nodes with spare CPU and memory, because the scheduler cannot see the constraint, and the symptom that does surface is an ENI attachment timeout rather than anything mentioning addresses [12]. Prefix delegation lifts one ENI prefix to 16 IPs instead of 1 [13], which is a change you make per cluster, not per chart.

The portable part of Kubernetes is the object you submit to the API server. Everything a process works out at runtime by reading the machine it landed on is a per-provider variable, and in this case a cheap one to check: `/proc/1/cgroup` returns a single `0::/` line on v2 and a `12:memory:/kubepods/...` line on v1 [9]. After about two years running all three, the author's conclusion was to stop treating cloud-agnostic as a real thing [4]. The narrower and more usable version: a chart is a claim about what the API server will accept, and the node image is the claim about how it will behave.

What to watch

  • Whether AKS and EKS node images follow GKE to cgroup v2, which would relocate the same failure to the clusters that currently pass.
  • Which other self-configuring runtimes in the estate read memory limits from the host rather than the container.
  • Whether node image versions get pinned per pool, or heap sizes get set explicitly, as the durable fix.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence42
Adoption
Insufficient
Hype gap+20
Incentives32
Confidence48
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    The same Helm chart, with identical values and the same image tag, was deployed to EKS, AKS and GKE within the same 20-minute release window.

    ReportedSupportedView cited source
  2. [2]

    EKS came up healthy and AKS came up healthy; GKE's pods started, ran for about 90 seconds, then began OOMKilling in a loop.

    ReportedSupportedView cited source
  3. [3]

    It took four hours to find why the GKE pods were being OOMKilled.

    ReportedSupportedView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. dev.to

    1 article · August 21, 2026

    I Managed Kubernetes Across AWS, Azure, and GCP Simultaneously — Here's What Nobody Tells You

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

  • cgroup v2 and container resource limitsFollow
  • Kubernetes workload identity pitfallsFollow
  • Cloud CNI and pod IP planningFollow
  • JVM and runtime container awarenessFollow
  • Kubernetes multi-cloud portabilityFollow
Loading related stories