LeadershipNot yet confirmed elsewhere1 publisher2 min readPublished
Kubernetes v1.35 kubelet refuses to start on cgroup v1 nodes by default
Kubernetes v1.35 sets failCgroupV1 to true by default, so the kubelet will not start on a node still running cgroup v1. Teams below v1.35 now decide before upgrading between migrating every Linux node and carrying a temporary override the project plans to remove.
The Board Room · Leadership desk

Clusters below v1.35 must migrate or plan the override. Clusters on v1.35 or later see leftover cgroup v1 nodes fail at kubelet startup. kubeadm preflight errors at init, join, upgrade. The override is temporary. I/O-heavy Pods still face evictions. Only cgroup v2 nodes get Memory QoS.
- decision Clusters below v1.35 Before upgrading, migrate every Linux node to cgroup v2 or plan the temporary failCgroupV1: false override, claim 4
- exposure Clusters on v1.35 or later Under the default configuration, any remaining cgroup v1 node fails during kubelet startup, claim 5
- constraint kubeadm-managed clusters SystemVerification preflight returns an error at init, join and upgrade on cgroup v1 with kubelet v1.35 or later, claim 6
- cost Teams keeping the override failCgroupV1: false is temporary; its removal follows the Kubernetes deprecation policy, claim 2
- exposure I/O-intensive workloads Moving to cgroup v2 does not change active_file counting; a large page cache can still trigger Pod evictions, claim 9
- capability Nodes migrated to cgroup v2 Only cgroup v2 nodes can run Memory QoS; cgroup v1 cannot provide its protection model, claim 12
| Who | How | Kind | Claim |
|---|---|---|---|
| Clusters below v1.35 | Before upgrading, migrate every Linux node to cgroup v2 or plan the temporary failCgroupV1: false override | decision | 4 |
| Clusters on v1.35 or later | Under the default configuration, any remaining cgroup v1 node fails during kubelet startup | exposure | 5 |
| kubeadm-managed clusters | SystemVerification preflight returns an error at init, join and upgrade on cgroup v1 with kubelet v1.35 or later | constraint | 6 |
| Teams keeping the override | failCgroupV1: false is temporary; its removal follows the Kubernetes deprecation policy | cost | 2 |
| I/O-intensive workloads | Moving to cgroup v2 does not change active_file counting; a large page cache can still trigger Pod evictions | exposure | 9 |
| Nodes migrated to cgroup v2 | Only cgroup v2 nodes can run Memory QoS; cgroup v1 cannot provide its protection model | capability | 12 |
What happened
- In kubeadm clusters, the v1.35 SystemVerification preflight check returns an error at init, join and upgrade when it finds cgroup v1 with a v1.35 or later kubelet.
- Moving to cgroup v2 does not change how the kubelet counts active_file memory, so large page caches can still trigger memory pressure and Pod evictions.
- Memory QoS, which remains alpha in v1.36, runs only on cgroup v2 nodes because cgroup v1 cannot provide its memory protection model.
Why it matters
- decision Choosing the override this quarter commits a team to the same migration at a later upgrade, on a timetable set by the Kubernetes deprecation policy instead of its own plan.
- constraint kubeadm fleets need a complete per-node cgroup inventory before the upgrade window opens, since the preflight check catches a missed node at the start of the operation.
- capability Migrated nodes can later adopt tiered memory protection for Guaranteed and Burstable Pods, an option cgroup v1 nodes will never have.
The default flip puts the failure at kubelet startup. On v1.35 or later, under the default configuration, any remaining cgroup v1 node fails as the kubelet starts [5]. For clusters still below v1.35, the project tells operators to migrate every Linux node to cgroup v2 before upgrading, or to plan for the temporary override [4]. Clusters already on v1.35 should confirm that every Linux node runs cgroup v2, or that any override is there on purpose [5].
The project gave a long notice period before the flip. By v1.35, cgroup v2 support had been stable for ten minor releases [18], and cgroup v1 had spent four releases in maintenance mode [19]. Kubernetes has deprecated cgroup v1 [16].
A platform lead with a full roadmap could fairly say the override makes this next year's problem. For the kubelet, that holds for a while. Operators can set failCgroupV1: false in the kubelet configuration file, and the project describes the setting as temporary [2]. Its removal will follow the Kubernetes deprecation policy, with the remaining work tracked in KEP-5573 [2][3]. We think the override is defensible for a team that needs v1.35 this quarter for other reasons and already has a dated migration plan.
kubeadm fleets hit the check sooner. The project calls the v1.35 SystemVerification preflight check an earlier, stricter check, and it applies to init, join and upgrade alike [6]. The post does not say whether the kubelet override changes that preflight result.
Teams hoping the move would end page-cache evictions on I/O-heavy workloads will still need the existing workaround. That means equal memory requests and limits for containers doing intensive I/O, set after measuring an appropriate value [10]. The underlying kubelet behaviour is tracked as kubernetes/kubernetes#43916 [9].
We would not count Memory QoS as a reason to migrate this quarter. The feature was introduced as alpha in v1.22 and is still alpha in v1.36 [11], fourteen minor releases later [20]. The project's general advice is to keep alpha features out of production [13]. Its new tiered mode maps Guaranteed Pod memory requests to hard protection and Burstable requests to soft protection [15]. Teams that enable it in production are told to test the configuration and account for hard-reserved memory first [17]. The recommendation is kernel 5.9 or later. Below that, a known livelock can be set off by memory.high reclaim, and starting with v1.36 the kubelet writes a warning to its log if Memory QoS is turned on with one of those affected kernels [14].
What to watch
- Which Kubernetes release removes the failCgroupV1 override under KEP-5573 and the deprecation policy.
- Project guidance on whether setting failCgroupV1: false also clears the kubeadm SystemVerification preflight error.
- Whether Memory QoS with tiered reservation leaves alpha in a release after v1.36.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence72
- Adoption
- Insufficient
- Hype gap0
- Incentives
- Insufficient
- Confidence68
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Starting with Kubernetes v1.35, failCgroupV1 defaults to true, so the kubelet does not start on a cgroup v1 node by default.
- [2]
Administrators can temporarily set failCgroupV1: false in the kubelet configuration file, but removal will follow the Kubernetes deprecation policy.
- [3]
Further removal work is tracked in KEP-5573: Remove cgroup v1 support.
- [4]
If you are still on a release older than v1.35, migrate every Linux node to cgroup v2 before upgrading, or plan to set the temporary failCgroupV1: false override.
- [5]
If already on v1.35 or later, confirm that every Linux node runs cgroup v2 (or that you intentionally keep the override). Under the default configuration, a remaining cgroup v1 node fails during kubelet startup.
- [6]
For kubeadm-managed clusters, Kubernetes v1.35 makes this an earlier, stricter check: the SystemVerification preflight check, provided by k8s.io/system-validators, returns an error during kubeadm init, kubeadm join and kubeadm upgrade when it detects cgroup v1 with kubelet v1.35 or later; with an older kubelet, the check remains a warning.
- [7]
Support for v2 cgroup management has been stable since Kubernetes v1.25.
- [8]
With the release of Kubernetes v1.31, support for v1 cgroup management moved into maintenance mode.
- [9]
The kubelet treats active_file memory as not reclaimable; for I/O-intensive workloads a large page cache can make the kubelet report memory pressure and evict Pods. This is a known kubelet issue (kubernetes/kubernetes#43916), and migrating to cgroup v2 does not by itself change that calculation.
- [10]
The documented workaround is to set equal memory requests and limits for containers that perform intensive I/O, after measuring an appropriate value.
- [11]
Memory QoS was introduced as an alpha feature in Kubernetes v1.22 and updated in v1.27; it remains alpha in v1.36, but now separates memory throttling from memory reservation and adds tiered memory protection.
- [12]
Memory QoS is available only on Linux nodes that use cgroup v2; cgroup v1 cannot provide this protection model.
- [13]
The overall Kubernetes recommendation is not to enable Alpha features in production.
- [14]
Kernel 5.9 or later is recommended for Memory QoS; on older kernels, memory.high reclaim can trigger a known livelock, and from v1.36 the kubelet logs a warning when Memory QoS is enabled on an affected kernel.
- [15]
memoryReservationPolicy: TieredReservation maps Guaranteed Pod memory requests to memory.min (hard protection) and Burstable Pod requests to memory.low (soft protection); BestEffort Pods receive neither.
- [17]
If using Memory QoS with tiered reservations in production, test the configuration and account for hard-reserved memory before enabling it.
- [18]
cgroup v2 support had been stable for ten minor releases by the time v1.35 made cgroup v1 nodes fail by default.
- [19]
cgroup v1 spent four minor releases in maintenance mode before v1.35 made it fail by default.
- [20]
Memory QoS has remained alpha across fourteen minor releases, from v1.22 to v1.36.
Sources
1 independent publisher whose own reporting we read for this story.
- kubernetes.ioThe Shift to cgroup v2 in Kubernetes: What You Need to Know
1 article · October 10, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.