BuildNot yet confirmed elsewhere1 publisher2 min readPublished
containerd 2.1 hands sandbox pods a writable cgroup without the privileged flag
containerd 2.1.0, shipped December 2024, added a cgroup_writable handler that gives a Kubernetes pod a writable cgroup subtree without privileged: true. Set hostUsers: false and the pod can create child cgroups but cannot raise its own memory limit.
The Engineer · Build desk
What happened
- The flag only ever bought a writable cgroup: the CRI mounts /sys/fs/cgroup read-only in every unprivileged container, so the one switch that makes it writable also hands the pod the rest of the node.
- Kubernetes already has pod fields for three of a sandbox's four needs - unconfined seccomp, an unmasked /proc and no AppArmor - but none for a writable cgroup v2 subtree.
- On the test cluster, mkdir /sys/fs/cgroup/child succeeds inside the pod while echo max > memory.max comes back Permission denied.
- Networked sandboxes hit a wall at /dev/net/tun, because a hostPath volume will not mount in a user-namespaced pod and the device node refuses an idmapped mount.
Why it matters
- exposure A sandbox builder that normally runs privileged: true can run without it, so a breakout inherits an unprivileged uid's own cgroup subtree instead of node-level access.
- cost This is a node-level change: operators add a containerd runtime drop-in and a RuntimeClass per node, so cluster teams own the change, not the people shipping pods.
- constraint The kubelet limit on the pod stays the hard ceiling, since memory.max remains owned by host root, so untrusted code in a sandbox cannot grow its budget past what the operator set.
On its own, a read-write /sys/fs/cgroup is more dangerous than the read-only one it replaces: a container running as node root could write max into its own memory.max and step past its budget. [5] What keeps the containerd handler safe is runc. [6]
If a container gets both a private cgroup namespace and a writable cgroupfs, runc hands ownership of its cgroup to whichever host uid the container's uid is mapped onto, then chowns the delegate files - cgroup.procs, cgroup.subtree_control, memory.oom.group and the rest - to that same uid. memory.max keeps root as its owner. [6] Put the pod in its own user namespace with hostUsers: false and that owner is an unprivileged id. [7] systemd calls this Delegate=yes; here it is scoped to one pod. [8]
The author ran it on k3s 1.36.5, containerd 2.3.4, runc 1.4.2 and Linux 6.8. [9] Without the handler, a user-namespaced pod sees the mount read-only and owned by 65534, the host's root left unmapped. [10] On a RuntimeClass that carries the handler, the same pod gets the mount read-write and owned by its own root, while memory.max still reports 65534. [11]
Making this work takes more than the pod spec. The node needs a containerd drop-in that defines a runtime with cgroup_writable = true and SystemdCgroup = true, plus a RuntimeClass that points the pod at that handler. [14] runAsUser: 0 is root of the pod's user namespace only, so runc chowns the cgroup to whatever uid the container starts as, and mapping ids into each sandbox needs CAP_SETUID there. [15]
For networked sandboxes, the /dev/net/tun problem has a fix: the handler's base_runtime_spec adds the device as containerd's default OCI spec plus the device node, generated on the node. [17] A CI script runs the whole setup against a fresh k3s and checks that the pod is unprivileged and that a sandbox actually runs. [18]
What to watch
- Whether Kubernetes adds a first-class pod field for a writable cgroup subtree, removing the per-node RuntimeClass and containerd drop-in step.
- Whether the runc chown behaviour the design leans on holds in runc releases after the 1.4.2 the author tested.
- How the setup behaves on managed clusters where operators cannot edit containerd config or register a RuntimeClass.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence60
- Adoption33
- Hype gap−5
- Incentives38
- Confidence55
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
A pod that builds sandboxes - one that runs other people's code in its own namespaces, with a cgroup per request - usually runs as privileged: true.
- [2]
A sandbox needs four things from the surrounding pod: user namespaces, a fully visible /proc, no AppArmor profile, and a cgroup v2 subtree it can write. Kubernetes has a pod field for the first three (seccompProfile: Unconfined, procMount: Unmasked, appArmorProfile: Unconfined) but no field for the writable cgroup subtree.
- [3]
The CRI mounts /sys/fs/cgroup read-only in every container that is not privileged. That is why privileged: true is used to get a writable cgroup, and it gives the pod far more than a cgroup.
- [4]
In December 2024 containerd added a runtime-handler option called cgroup_writable, shipped in v2.1.0. For containers started through that handler, /sys/fs/cgroup is mounted read-write.
- [5]
On its own, a read-write cgroupfs would be dangerous: a container that is root on the node could write max into its own memory.max.
- [6]
What makes it safe is runc. When the container has its own cgroup namespace and a read-write cgroupfs, runc sets the cgroup's owner to the host uid that the container's uid maps to, and its systemd cgroup driver chowns the cgroup directory and the delegate files - cgroup.procs, cgroup.threads, cgroup.subtree_control, memory.oom.group, memory.reclaim - to that owner. memory.max stays root's.
- [7]
In a pod with hostUsers: false, that cgroup owner is an unprivileged uid.
- [8]
The pod can create cgroups below its own but cannot raise its own limit. This is systemd's Delegate=yes, for a pod.
- [9]
The setup was tested on k3s 1.36.5, containerd 2.3.4, runc 1.4.2 and Linux 6.8.
- [10]
A user-namespaced pod without the handler sees /sys/fs/cgroup mounted read-only, owner 65534 (the host's root, unmapped).
- [11]
The same pod on a RuntimeClass with the handler sees /sys/fs/cgroup read-write owned by uid 0 (the pod's root), cgroup.procs owned by 0, and memory.max owned by 65534.
- [12]
Inside the pod, mkdir /sys/fs/cgroup/child succeeds while echo max > /sys/fs/cgroup/memory.max returns Permission denied.
- [14]
The setup needs a containerd drop-in on the node that defines a runtime with runtime_type io.containerd.runc.v2, cgroup_writable = true and SystemdCgroup = true, plus a RuntimeClass that points the pod at that handler.
- [15]
runAsUser: 0 is root of the pod's user namespace only. runc chowns the cgroup to the uid the container starts as, and mapping ids into each sandbox's namespace needs CAP_SETUID there.
- [16]
Networked sandboxes need /dev/net/tun, but a hostPath volume does not start in a user-namespaced pod: the kubelet asks for an idmapped mount and the device node refuses it (failed to set MOUNT_ATTR_IDMAP on /dev/net/tun).
- [17]
The handler's base_runtime_spec adds the device instead - containerd's default OCI spec plus the device node, generated on the node.
- [18]
A script runs the setup against a fresh k3s in CI and checks that the pod is not privileged and that a sandbox runs.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toSandboxes in Kubernetes without privileged: cgroup_writable and hostUsers: false
1 article · October 6, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Linux cgroupsFollow
- Container SandboxingFollow
- Kubernetes SecurityFollow
- Container runtimesFollow