LeadershipNot yet confirmed elsewhere1 publisher3 min readPublished
Kubernetes ties node swap's up-to-3x density gain to memory that sits idle
Kubernetes' project blog says NVMe-backed node swap, generally available since v1.34, packed up to 3x more pods onto memory-bound nodes in benchmarks. The gain comes from paging out memory that agent sandboxes hold idle between prompts, so swap-based budgets depend on how much of each pod is truly dormant.
The Board Room · Leadership desk

Bar comparison of container memory limits for a kernel build. Without swap, 600 MB was the minimum to avoid an OOM crash. With swap on a Local SSD, 300 MB ran with no slowdown, 374s versus the 433s baseline. At 200 MB the active working set went into swap and run time rose by over 40%.
Container memory limit, Linux kernel build benchmark
| Item | Value | Claim |
|---|---|---|
| No swap: minimum to avoid OOM | 600 MB | 3 |
| Swap on SSD: no slowdown | 300 MB | 4 |
| Swap on SSD: over 40% slower | 200 MB | 5 |
What happened
- The benchmarks covered three workloads: CI/CD kernel builds, sandboxed headless browsers and isolated Python runtimes.
- At a 200 MB limit the build's active working set went into swap, and run time rose by more than 40%.
- According to the post, Kubernetes swap support relies on cgroup v2's separate swap accounting in place of cgroup v1's single combined limit.
Why it matters
- constraint Cutting limits below a pod's active working set turns swap into a slowdown, and in the kernel build that point arrived at the setting that would have tripled packing.
- decision Capacity plans for agent sandboxes have to rest on measured idle memory per workload, since the up-to-3x ceiling is a best case summarised across three different jobs.
- cost The RAM saving assumes every node carries fast local NVMe and runs cgroup v2, so drive and OS standards become part of the memory budget.
The kernel-build test shows where the density gain stops. Without swap, a Linux 6.1.1 build needed a memory limit of at least 600 MB to avoid an out-of-memory kill [3]. With swap on a local SSD, the limit fell to 300 MB and the build finished in 374 seconds against 433 [4], roughly 14% faster [13]. At 200 MB it ran more than 40% slower, with long I/O waits, because its active working set had been forced into swap [5]. A 200 MB limit is one third of the baseline. On memory alone, that is the setting that fits three builds in the RAM one needed without swap [12].
The gap between 300 MB and 200 MB comes from what a build keeps in memory. Earlier compiled objects sit inactive while the pipeline moves on, and the linking phase needs a brief, large spike [6]. The inactive objects paged out with no cost to run time. Once the limit cut into memory the build was actively using, the I/O waits began [4][5]. The post's authors wrote that the result shows "swap serves as an insurance policy for burst memory, not a replacement for active RAM." [10]
Agent sandboxes look like the inactive half of that build. The post describes agent pods that need large footprints to start and to run untrusted code, then sit idle for long stretches waiting for user prompts [2]. That resident, idle memory is what swap moves to disk. The headline figure, density gains of up to 3x "often with little or no latency cost", is a summary across all three workloads [14][15]. The summary does not say which workload reached 3x.
A skeptic would point out that Kubernetes historically discouraged swap, and the post agrees there were reasons. Under cgroup v1, memory and swap shared one combined limit, so a container's real memory use was unpredictable and hard to isolate [7]. The GA support relies on cgroup v2, which accounts for swap separately [8]. Paging to spinning disks was slow, a penalty the post says fast NVMe local SSDs largely eliminate [9]. Both answers are conditions. A node without cgroup v2 or without fast local drives is still subject to the old objections [7][9].
In our view, memory budgets can be planned around swap for workloads whose idle share has been measured, with limits kept at or above the active working set. Planning a whole fleet around the 3x ceiling goes further than the evidence. The benchmarks are the project's own, published on its blog [11]. The decision this quarter is where to set limits. Next quarter's consequence is how packed nodes behave when more sandboxes are active at once than the plan assumed. For one workload, the post has already measured what active memory in swap costs: a run more than 40% slower [5].
What to watch
- Whether the browser and Python sandbox results show the same slowdown once limits reach the active working set.
- Independent replication of the up-to-3x density figure on production agent fleets, outside the project's own benchmarks.
- Latency data for many idle sandboxes resuming from swap at once, the load a tightly packed node would face.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives60
- Confidence50
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Kubernetes support for running nodes with swap enabled reached General Availability in v1.34.
- [2]
Agentic pods require large memory footprints to initialize and execute untrusted code, then typically enter long-tail idle phases waiting for user prompts; keeping this idle state in physical RAM caps cluster density.
- [3]
On a baseline node without swap, the minimum memory limit to prevent an OOM crash during a Linux 6.1.1 kernel build was 600 MB.
- [4]
Routing swap to a Local SSD cut the container memory limit by 50% to 300 MB without any execution slowdown; the build ran in 374s versus the baseline 433s.
- [5]
Compressing the limit to 200 MB forced the active working set into swap, causing long I/O wait times and increasing execution time by over 40%.
- [6]
In CI/CD jobs, earlier compiled objects sit inactive in memory while the pipeline progresses; the kernel build requires a large memory spike during the brief linking phase.
- [7]
Under cgroup v1, memory and swap were treated as a single combined limit, which made a container's real memory usage unpredictable and hard to isolate; this was one reason swap was historically discouraged in Kubernetes.
- [8]
Kubernetes' swap support relies on cgroup v2, whose separate swap accounting tracks disk swap on its own.
- [9]
The second historical reason swap was discouraged was the latency penalty of paging to slow spinning disks, which fast NVMe Local SSDs largely eliminate.
- [10]
"swap serves as an insurance policy for burst memory, not a replacement for active RAM."
- [11]
The benchmarks were run by the post's authors and published on the Kubernetes project blog.
- [12]
A 200 MB limit is one third of the 600 MB no-swap baseline, equivalent on memory alone to fitting three builds in the RAM one build needed without swap.
- [13]
The 300 MB swap-backed build ran roughly 14% faster than the baseline.
- [14]
By backing node swap with fast NVMe SSDs, a node can page out dormant memory and pack in far more pods; the benchmarks found density gains of up to 3x, often with little or no latency cost.
ReportedInsufficientSource: Kubernetes project blog2 sources— create a free account to open themView cited source - [15]
The benchmarks covered three workloads: CI/CD kernel builds, sandboxed headless browsers, and isolated Python runtimes.
ReportedInsufficientSource: Kubernetes project blog2 sources— create a free account to open themView cited source
Sources
1 independent publisher whose own reporting we read for this story.
- kubernetes.ioScaling Kubernetes Workloads with Node Swap
1 article · October 10, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Memory overcommit and pod densityFollow
- AI Agent SandboxingFollow
- Kubernetes node swapFollow
Entities
- KubernetesFollow
- cgroup v2Follow
- gVisorFollow
- Kata ContainersFollow
- agent-sandboxFollow
- runcFollow