BuildNot yet confirmed elsewhere1 publisher3 min readPublished
Exit code 137 is the kernel collecting on a bet you did not know it placed
Kubernetes calls it OOMKilled. Linux calls it overcommit coming due, and a badness score nobody in the deploy pipeline has looked at decides who pays.
The Engineer · Build desk
What happened
- An OOM kill reaches Kubernetes as exit code 137 and OOMKilled: a kernel SIGKILL with no exception, no core dump and no shutdown hook.
- Linux allocates optimistically by default, returning address space from malloc() without checking that physical memory exists to back it.
- Host kills are logged in dmesg; cgroup kills show up in journalctl -k and in the oom_kill counter in memory.events.
Why it matters
- constraint Nothing in the process gets to observe its own death, so the evidence has to be pulled from kernel counters rather than application logging or a handler you were planning to add.
- decision The oom_kill counter against the node memory graph decides whether you are editing a manifest or resizing hardware, and that branch is taken before any code is read.
- exposure Any critical process that has never been killed may owe that only to a fatter neighbour rather than to a set adjustment, and the next deployment can reorder the queue.
- cost Misreading the signal is paid in on-call hours spent chasing a leak in the wrong subsystem, with the pod restarting the whole time.
The useful question after a 137 is not why the application asked for so much memory. It is which of the two killers fired, because they have different triggers and neither fix lives in the application repo [12][13].
Start with the bet. When a process calls `malloc()`, the kernel usually hands back a virtual address range without checking that physical memory exists to back it, a behaviour called overcommit and controlled by `vm.overcommit_memory`, which defaults to heuristic mode 0 [4][5]. The reason is `fork()`: a process holding 4GB of RSS that forks would, under strict accounting, have to reserve a second 4GB it will almost certainly never touch because of copy-on-write [6]. That is 8GB booked to run one 4GB workload [15], and refusing the fork would break an enormous amount of software that relies on cheap forking [6]. So the kernel defers the reckoning until pages are actually touched rather than requested [7].
The deferral is why there is nothing to catch. When the kernel genuinely cannot satisfy a page fault, something has to die immediately and synchronously inside the page allocation path [7], which is why the process gets a SIGKILL with no exception, no core dump and no shutdown hook [2], and why the only trace on a host-level kill is one line in `dmesg` [3].
Victim selection is the part that decides how long the incident runs. The kernel does not kill the process whose allocation pushed the system over the edge; it kills whichever process scores worst on a badness heuristic, readable at `/proc/<pid>/oom_score` [8]. That score is roughly resident memory plus swap as a share of total system memory, then adjusted by `oom_score_adj`, a value from -1000 to 1000 that defaults to 0 [9]. A process pinned at -1000 is treated as unkillable and survives even while using 90% of RAM, which is how `sshd` and `systemd` stay up so you can log in afterwards [10]. Read that in reverse: anything in your critical path that has never been killed has been surviving because it happened to be smaller than a noisier neighbour, unless somebody set the adjustment on purpose [9][10].
Inside a container the denominator changes. Kubernetes writes a memory limit as `memory.max` in the pod's cgroup v2 hierarchy, and when the cgroup's usage crosses that ceiling the kernel runs an OOM kill scoped to processes in that cgroup alone [11]. Host memory is irrelevant to that decision, which is why a pod dies while `free -h` on the node looks comfortable [12]. On bare metal or a plain VPS the victim pool is every process on the box instead [14].
Hence the one diagnostic that matters. Host kills surface through `dmesg -T | grep -i 'killed process'`; cgroup kills surface through `journalctl -k` and the `oom_kill` counter in `memory.events` under the cgroup path [13]. A climbing `oom_kill` counter alongside a flat node memory graph is a limit set too tight, not a leak [1]. Check the counter first and the fix is a number in a manifest; skip it and you are reading heap profiles for an allocation that was never the problem, which is the gap the dev.to write-up puts at five minutes versus three hours [16].
What to watch
- Pods already closed out as application crashes that turn out to have a climbing oom_kill counter in memory.events.
- Kills that show up in dmesg rather than a cgroup counter, which move the fix from pod limits to node sizing and a system-wide victim pool.
- Whether anything in the critical path carries a deliberate non-zero oom_score_adj, or whether survival has been accidental.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+8
- Incentives25
- Confidence55
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
A rising oom_kill counter with a stable node-level memory graph is the signature of a cgroup limit that is too tight, not a real leak.
- [2]
A kernel OOM kill appears in Kubernetes as exit code 137 with status OOMKilled: a SIGKILL from the kernel with no stack trace, no exception, no core dump, no warning and no graceful shutdown hook.
- [3]
A host-level kill leaves a single dmesg line of the form 'Out of memory: Killed process 4821 (node)'.
- [4]
Linux's default allocation strategy is optimistic: when a process calls malloc(), the kernel usually returns a virtual address range immediately without checking whether enough physical memory exists to back it. This is overcommit, controlled by vm.overcommit_memory.
- [5]
vm.overcommit_memory values: 0 is heuristic overcommit and the default, 1 always overcommits and never refuses an allocation, 2 is strict accounting that refuses once the commit limit is hit.
- [6]
Overcommit exists because of fork(): a process using 4GB of RSS that forks would, under strict accounting, need to reserve another 4GB it will almost certainly never touch thanks to copy-on-write, and refusing that fork would break enormous amounts of software that relies on cheap forking.
- [7]
Default heuristic mode defers the reckoning until pages are touched rather than requested; when the kernel genuinely cannot satisfy a page fault, something has to die immediately and synchronously inside the kernel's page allocation path.
- [8]
The OOM killer does not kill the process whose allocation pushed the system over the edge; it kills whichever process scores worst on a badness heuristic, which can be read at /proc/<pid>/oom_score.
- [9]
The badness calculation is roughly the process's resident memory plus swap usage as a percentage of total system memory, then adjusted by oom_score_adj, a per-process value from -1000 to 1000 that defaults to 0 and can be set by a user or process supervisor.
- [10]
A process set to oom_score_adj -1000 is exempted and treated as unkillable, surviving even if it is using 90% of RAM; sshd and systemd typically do this, which is why sshd stays alive while an app server dies.
- [11]
Kubernetes writes a container memory limit as memory.max in the pod's cgroup v2 hierarchy; when the cgroup's usage exceeds that ceiling the kernel invokes an OOM killer scoped to that cgroup and picks a victim from the processes inside it.
- [12]
The cgroup memory controller's OOM kill is a completely separate trigger from the system-wide OOM killer, which is why a single pod can be OOMKilled while the node still shows plenty of free memory in free -h.
- [13]
Host-level kills are visible via dmesg -T | grep -i 'killed process'; cgroup-level kills inside a container are found via journalctl -k | grep -i oom and the oom_kill counter in /sys/fs/cgroup/<path>/memory.events.
- [14]
On bare metal or a plain VPS, the OOM killer picks its victim from every process on the system.
- [15]
Under strict accounting, the 4GB forking process in the source's example would require 8GB of commit to run a 4GB workload.
- [16]
The dev.to write-up argues that understanding how the OOM killer decides who dies is the difference between a five-minute fix and a three-hour production incident, and that OOM kills should not be treated as random.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toWhat Actually Happens When Linux Runs Out of Memory: Inside the OOM Killer
1 article · August 23, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Linux Memory ManagementFollow
- Kubernetes OperationsFollow
- Production Incident DebuggingFollow
- Container Resource LimitsFollow