Skip to content

LeadershipNot yet confirmed elsewhere1 publisher3 min readPublished

Kubernetes ties node swap's up-to-3x density gain to memory that sits idle

Kubernetes' project blog says NVMe-backed node swap, generally available since v1.34, packed up to 3x more pods onto memory-bound nodes in benchmarks. The gain comes from paging out memory that agent sandboxes hold idle between prompts, so swap-based budgets depend on how much of each pod is truly dormant.

The Board Room · Leadership desk

How we use AISend a correction

Illustration accompanying Kubernetes ties node swap's up-to-3x density gain to memory that sits idle
Generated illustration
Kernel build held up at 300 MB, slowed at 200 MB Container memory limit in the authors' Linux kernel build benchmark: the 600 MB no-swap minimum versus two limits backed by swap on a Local SSD.

Bar comparison of container memory limits for a kernel build. Without swap, 600 MB was the minimum to avoid an OOM crash. With swap on a Local SSD, 300 MB ran with no slowdown, 374s versus the 433s baseline. At 200 MB the active working set went into swap and run time rose by over 40%.

Kernel build held up at 300 MB, slowed at 200 MB (Container memory limit, Linux kernel build benchmark)
ItemValueClaim
No swap: minimum to avoid OOM600 MB3
Swap on SSD: no slowdown300 MB4
Swap on SSD: over 40% slower200 MB5

What happened

  • The benchmarks covered three workloads: CI/CD kernel builds, sandboxed headless browsers and isolated Python runtimes.
  • At a 200 MB limit the build's active working set went into swap, and run time rose by more than 40%.
  • According to the post, Kubernetes swap support relies on cgroup v2's separate swap accounting in place of cgroup v1's single combined limit.

Why it matters

  • constraint Cutting limits below a pod's active working set turns swap into a slowdown, and in the kernel build that point arrived at the setting that would have tripled packing.
  • decision Capacity plans for agent sandboxes have to rest on measured idle memory per workload, since the up-to-3x ceiling is a best case summarised across three different jobs.
  • cost The RAM saving assumes every node carries fast local NVMe and runs cgroup v2, so drive and OS standards become part of the memory budget.

The kernel-build test shows where the density gain stops. Without swap, a Linux 6.1.1 build needed a memory limit of at least 600 MB to avoid an out-of-memory kill [3]. With swap on a local SSD, the limit fell to 300 MB and the build finished in 374 seconds against 433 [4], roughly 14% faster [13]. At 200 MB it ran more than 40% slower, with long I/O waits, because its active working set had been forced into swap [5]. A 200 MB limit is one third of the baseline. On memory alone, that is the setting that fits three builds in the RAM one needed without swap [12].

The gap between 300 MB and 200 MB comes from what a build keeps in memory. Earlier compiled objects sit inactive while the pipeline moves on, and the linking phase needs a brief, large spike [6]. The inactive objects paged out with no cost to run time. Once the limit cut into memory the build was actively using, the I/O waits began [4][5]. The post's authors wrote that the result shows "swap serves as an insurance policy for burst memory, not a replacement for active RAM." [10]

Agent sandboxes look like the inactive half of that build. The post describes agent pods that need large footprints to start and to run untrusted code, then sit idle for long stretches waiting for user prompts [2]. That resident, idle memory is what swap moves to disk. The headline figure, density gains of up to 3x "often with little or no latency cost", is a summary across all three workloads [14][15]. The summary does not say which workload reached 3x.

A skeptic would point out that Kubernetes historically discouraged swap, and the post agrees there were reasons. Under cgroup v1, memory and swap shared one combined limit, so a container's real memory use was unpredictable and hard to isolate [7]. The GA support relies on cgroup v2, which accounts for swap separately [8]. Paging to spinning disks was slow, a penalty the post says fast NVMe local SSDs largely eliminate [9]. Both answers are conditions. A node without cgroup v2 or without fast local drives is still subject to the old objections [7][9].

In our view, memory budgets can be planned around swap for workloads whose idle share has been measured, with limits kept at or above the active working set. Planning a whole fleet around the 3x ceiling goes further than the evidence. The benchmarks are the project's own, published on its blog [11]. The decision this quarter is where to set limits. Next quarter's consequence is how packed nodes behave when more sandboxes are active at once than the plan assumed. For one workload, the post has already measured what active memory in swap costs: a run more than 40% slower [5].

What to watch

  • Whether the browser and Python sandbox results show the same slowdown once limits reach the active working set.
  • Independent replication of the up-to-3x density figure on production agent fleets, outside the project's own benchmarks.
  • Latency data for many idle sandboxes resuming from swap at once, the load a tightly packed node would face.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence45
Adoption
Insufficient
Hype gap+25
Incentives60
Confidence50
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Kubernetes support for running nodes with swap enabled reached General Availability in v1.34.

    ReportedSupportedSource: Kubernetes project blogView cited source
  2. [2]

    Agentic pods require large memory footprints to initialize and execute untrusted code, then typically enter long-tail idle phases waiting for user prompts; keeping this idle state in physical RAM caps cluster density.

    ReportedSupportedSource: Kubernetes project blogView cited source
  3. [3]

    On a baseline node without swap, the minimum memory limit to prevent an OOM crash during a Linux 6.1.1 kernel build was 600 MB.

    ReportedSupportedSource: Kubernetes project blogView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. kubernetes.io

    1 article · October 10, 2026

    Scaling Kubernetes Workloads with Node Swap

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Loading related stories