Build1 publisher3 min readPublished
Container checkpoint restore skips the step that turns a Pod's security context into kernel state
containerd's September 1 advisory says a container restored from an untrusted checkpoint can run as root with full capabilities despite a restrictive Pod spec. Admission approves the spec, and restore then replays saved state without the step that turns a spec into kernel settings.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- The containerd advisory applies where checkpoint restore through CRI is enabled and an attacker can run a container from a crafted checkpoint image.
- Affected containerd versions run from 2.1.0 to before 2.2.7 and from 2.3.0 to before 2.3.4.
- CRI-O's CVE-2026-92574 records the same outcome, with saved credentials, capabilities, no_new_privs and seccomp state taking effect over the destination's configuration.
- Google's GKE security bulletins list three earlier trust failures on the same checkpoint and restore path in June.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint Tightening Pod security policy at admission does not reach a workload started from a checkpoint image, because the spec that passes review is not the one the runtime applies.
- exposure Anyone allowed to create Pods on a node with restore enabled can bring up a process carrying the privileges saved in an image, past the security context the Pod declares.
- decision With restore off by default in the fixed releases, re-enabling it becomes an explicit choice to run a path the advisories show can override the Pod spec.
- precedent If the author's frame holds, fixes that patch one attribute at a time will keep producing bulletins until runtimes define which side wins when saved state and policy disagree.
On the normal path a process is built from declarations. A Pod spec passes admission, the kubelet hands the runtime a container config, and the runtime turns that config into kernel state: a user and group, a capability set, the no_new_privs flag, a seccomp filter [2]. Kubernetes documents that allowPrivilegeEscalation directly controls whether no_new_privs is set [3]. So the field in the spec is a request, and the flag on the running process is what the kernel enforces. Only the runtime's translation connects the two [2].
Restore skips that step. CRIU records a running process, including its credentials, capabilities, no_new_privs flag and seccomp state, and restore replays that record [4]. CRIU is doing what it was built for: a process saved as root comes back as root [4].
The harder constraint sits upstream of the runtime. On the path containerd exposed, restore is triggered inside the ordinary container-create call, by a checkpoint archive or an annotated image [5]. Kubernetes supports container restore only through those image annotations, so admission sees a Pod create with an image reference [6]. The runtime decides later whether it is a restore, based on what the image contains [6]. Any policy enforced at admission is therefore checking a security context that, on this path, never becomes kernel state [1]. In the containerd case the restored container also ran with no seccomp filter, despite the restrictive policy the orchestrator requested [7].
The dev.to post that collected these advisories argues they are one event: state supplied by an artifact took effect on restore, and the destination's policy was not applied over it [14]. By its account, the runtime did not weigh the saved credentials against the requested user and pick the saved ones. The translation that would have replaced them never ran [15]. "A precedence rule that was chosen can be audited and changed. One that was never made has to be built," the author wrote [16]. I think the author is right. The June failures fit the pattern, because each one accepted something the checkpoint supplied [10].
Exposure is gated. Each of these issues needs the ability to create Pods, and the September one matters only where checkpoint restore is enabled [12]. Google pairs its Medium rating for GKE with the note that default GKE nodes ship without the criu binary [13]. For that rating to transfer to another cluster, its nodes would need the same property, with restore through CRI left off [8]. A self-managed node with criu installed and the feature switched on is closer to the case containerd rated Critical [13].
The fixed releases close the route by disabling it by default [17]. The post does not say whether those releases apply the Pod's security context over saved state when an operator turns restore back on.
What to watch
- Whether containerd or CRI-O adds a rule that applies the Pod's security context over saved credentials, capabilities and seccomp state when restore is enabled.
- Whether Kubernetes exposes restore to admission as its own operation instead of an image annotation the runtime interprets.
- Further GKE security bulletins on the checkpoint and restore path after the June and September issues.