Skip to content

BuildNot yet confirmed elsewhere1 publisher2 min readPublished

containerd 2.1 hands sandbox pods a writable cgroup without the privileged flag

containerd 2.1.0, shipped December 2024, added a cgroup_writable handler that gives a Kubernetes pod a writable cgroup subtree without privileged: true. Set hostUsers: false and the pod can create child cgroups but cannot raise its own memory limit.

The Engineer · Build desk

How we use AISend a correction

What happened

  • The flag only ever bought a writable cgroup: the CRI mounts /sys/fs/cgroup read-only in every unprivileged container, so the one switch that makes it writable also hands the pod the rest of the node.
  • Kubernetes already has pod fields for three of a sandbox's four needs - unconfined seccomp, an unmasked /proc and no AppArmor - but none for a writable cgroup v2 subtree.
  • On the test cluster, mkdir /sys/fs/cgroup/child succeeds inside the pod while echo max > memory.max comes back Permission denied.
  • Networked sandboxes hit a wall at /dev/net/tun, because a hostPath volume will not mount in a user-namespaced pod and the device node refuses an idmapped mount.

Why it matters

  • exposure A sandbox builder that normally runs privileged: true can run without it, so a breakout inherits an unprivileged uid's own cgroup subtree instead of node-level access.
  • cost This is a node-level change: operators add a containerd runtime drop-in and a RuntimeClass per node, so cluster teams own the change, not the people shipping pods.
  • constraint The kubelet limit on the pod stays the hard ceiling, since memory.max remains owned by host root, so untrusted code in a sandbox cannot grow its budget past what the operator set.

On its own, a read-write /sys/fs/cgroup is more dangerous than the read-only one it replaces: a container running as node root could write max into its own memory.max and step past its budget. [5] What keeps the containerd handler safe is runc. [6]

If a container gets both a private cgroup namespace and a writable cgroupfs, runc hands ownership of its cgroup to whichever host uid the container's uid is mapped onto, then chowns the delegate files - cgroup.procs, cgroup.subtree_control, memory.oom.group and the rest - to that same uid. memory.max keeps root as its owner. [6] Put the pod in its own user namespace with hostUsers: false and that owner is an unprivileged id. [7] systemd calls this Delegate=yes; here it is scoped to one pod. [8]

The author ran it on k3s 1.36.5, containerd 2.3.4, runc 1.4.2 and Linux 6.8. [9] Without the handler, a user-namespaced pod sees the mount read-only and owned by 65534, the host's root left unmapped. [10] On a RuntimeClass that carries the handler, the same pod gets the mount read-write and owned by its own root, while memory.max still reports 65534. [11]

Making this work takes more than the pod spec. The node needs a containerd drop-in that defines a runtime with cgroup_writable = true and SystemdCgroup = true, plus a RuntimeClass that points the pod at that handler. [14] runAsUser: 0 is root of the pod's user namespace only, so runc chowns the cgroup to whatever uid the container starts as, and mapping ids into each sandbox needs CAP_SETUID there. [15]

For networked sandboxes, the /dev/net/tun problem has a fix: the handler's base_runtime_spec adds the device as containerd's default OCI spec plus the device node, generated on the node. [17] A CI script runs the whole setup against a fresh k3s and checks that the pod is unprivileged and that a sandbox actually runs. [18]

What to watch

  • Whether Kubernetes adds a first-class pod field for a writable cgroup subtree, removing the per-node RuntimeClass and containerd drop-in step.
  • Whether the runc chown behaviour the design leans on holds in runc releases after the 1.4.2 the author tested.
  • How the setup behaves on managed clusters where operators cannot edit containerd config or register a RuntimeClass.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence60
Adoption33
Hype gap−5
Incentives38
Confidence55
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    A pod that builds sandboxes - one that runs other people's code in its own namespaces, with a cgroup per request - usually runs as privileged: true.

    ReportedSupportedView cited source
  2. [2]

    A sandbox needs four things from the surrounding pod: user namespaces, a fully visible /proc, no AppArmor profile, and a cgroup v2 subtree it can write. Kubernetes has a pod field for the first three (seccompProfile: Unconfined, procMount: Unmasked, appArmorProfile: Unconfined) but no field for the writable cgroup subtree.

    ReportedSupportedView cited source
  3. [3]

    The CRI mounts /sys/fs/cgroup read-only in every container that is not privileged. That is why privileged: true is used to get a writable cgroup, and it gives the pod far more than a cgroup.

    ReportedSupportedView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. dev.to

    1 article · October 6, 2026

    Sandboxes in Kubernetes without privileged: cgroup_writable and hostUsers: false

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories