Skip to content

Build1 publisher3 min readPublished

Container checkpoint restore skips the step that turns a Pod's security context into kernel state

containerd's September 1 advisory says a container restored from an untrusted checkpoint can run as root with full capabilities despite a restrictive Pod spec. Admission approves the spec, and restore then replays saved state without the step that turns a spec into kernel settings.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • The containerd advisory applies where checkpoint restore through CRI is enabled and an attacker can run a container from a crafted checkpoint image.
  • Affected containerd versions run from 2.1.0 to before 2.2.7 and from 2.3.0 to before 2.3.4.
  • CRI-O's CVE-2026-92574 records the same outcome, with saved credentials, capabilities, no_new_privs and seccomp state taking effect over the destination's configuration.
  • Google's GKE security bulletins list three earlier trust failures on the same checkpoint and restore path in June.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Tightening Pod security policy at admission does not reach a workload started from a checkpoint image, because the spec that passes review is not the one the runtime applies.
  • exposure Anyone allowed to create Pods on a node with restore enabled can bring up a process carrying the privileges saved in an image, past the security context the Pod declares.
  • decision With restore off by default in the fixed releases, re-enabling it becomes an explicit choice to run a path the advisories show can override the Pod spec.
  • precedent If the author's frame holds, fixes that patch one attribute at a time will keep producing bulletins until runtimes define which side wins when saved state and policy disagree.

On the normal path a process is built from declarations. A Pod spec passes admission, the kubelet hands the runtime a container config, and the runtime turns that config into kernel state: a user and group, a capability set, the no_new_privs flag, a seccomp filter [2]. Kubernetes documents that allowPrivilegeEscalation directly controls whether no_new_privs is set [3]. So the field in the spec is a request, and the flag on the running process is what the kernel enforces. Only the runtime's translation connects the two [2].

Restore skips that step. CRIU records a running process, including its credentials, capabilities, no_new_privs flag and seccomp state, and restore replays that record [4]. CRIU is doing what it was built for: a process saved as root comes back as root [4].

The harder constraint sits upstream of the runtime. On the path containerd exposed, restore is triggered inside the ordinary container-create call, by a checkpoint archive or an annotated image [5]. Kubernetes supports container restore only through those image annotations, so admission sees a Pod create with an image reference [6]. The runtime decides later whether it is a restore, based on what the image contains [6]. Any policy enforced at admission is therefore checking a security context that, on this path, never becomes kernel state [1]. In the containerd case the restored container also ran with no seccomp filter, despite the restrictive policy the orchestrator requested [7].

The dev.to post that collected these advisories argues they are one event: state supplied by an artifact took effect on restore, and the destination's policy was not applied over it [14]. By its account, the runtime did not weigh the saved credentials against the requested user and pick the saved ones. The translation that would have replaced them never ran [15]. "A precedence rule that was chosen can be audited and changed. One that was never made has to be built," the author wrote [16]. I think the author is right. The June failures fit the pattern, because each one accepted something the checkpoint supplied [10].

Exposure is gated. Each of these issues needs the ability to create Pods, and the September one matters only where checkpoint restore is enabled [12]. Google pairs its Medium rating for GKE with the note that default GKE nodes ship without the criu binary [13]. For that rating to transfer to another cluster, its nodes would need the same property, with restore through CRI left off [8]. A self-managed node with criu installed and the feature switched on is closer to the case containerd rated Critical [13].

The fixed releases close the route by disabling it by default [17]. The post does not say whether those releases apply the Pod's security context over saved state when an operator turns restore back on.

What to watch

  • Whether containerd or CRI-O adds a rule that applies the Pod's security context over saved credentials, capabilities and seccomp state when restore is enabled.
  • Whether Kubernetes exposes restore to admission as its own operation instead of an image annotation the runtime interprets.
  • Further GKE security bulletins on the checkpoint and restore path after the June and September issues.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories