Skip to content

Build1 publisher3 min readPublished

Laravel 13.34 lets queued jobs count OOM kills against their exception limit

Laravel 13.33 and 13.34 add opt-in guards for two of the three ways a queued job dies in Docker, worker timeouts and crashes. Neither touches Docker's stop grace period, so a slow model call still gets ten seconds by default after a deploy's SIGTERM.

The Engineer · Build desk

Illustration accompanying Laravel 13.34 lets queued jobs count OOM kills against their exception limit

What happened

  • Laravel 13.34.0 adds a per-job CountCrashesAsExceptions attribute, so an attempt that died without throwing counts as one exception toward $maxExceptions on the next attempt.
  • Without the attribute, a job that OOM-kills every worker it reaches loops until retryUntil() expires, because maxExceptions only counts exceptions that were actually thrown.
  • Laravel 13.33 adds a static Worker::$killOnTimeout flag, and when it is set to false the worker throws TimeoutExceededException inside the job instead of killing itself with SIGKILL.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Because crash counting is opt-in per job, any model-calling job class that nobody updates keeps paying for OOM loops in billed requests until its retryUntil() window closes.
  • constraint On the array driver, or a file cache inside separate containers, the attribute never finds a leftover marker, so crash-looping jobs stay uncounted even with the guard in place.
  • decision Switching killOnTimeout off is only safe once handle() methods are free of catch-all Throwable blocks and blocking system calls; otherwise a timeout can be swallowed or never delivered.

An OOM kill and a timeout end with the same signal, but only the timeout comes from Laravel itself [3][4]. On a timeout the worker uses pcntl and SIGALRM. When the alarm fires, it has always killed itself with SIGKILL. The timeout is recorded against the job, which fails at once if it sets `$failOnTimeout` [3]. An out-of-memory kill comes from the kernel when the process crosses the container's memory limit. There is no exception, no `failed()` callback and no application log, and if `php` is PID 1 the container exits with 137 [4]. A deploy comes from Docker. `docker compose up` sends SIGTERM, the worker tries to finish its current job, and SIGKILL follows after `stop_grace_period`, ten seconds by default [5].

After an OOM kill or a deploy kill, the job stays reserved until `retry_after` expires and then goes to another worker. The framework counts nothing, because nothing was thrown [6]. Jobs that call rate-limited APIs often cannot rely on `$tries`. According to the post, the `RateLimited` middleware releases the job and burns an attempt even when the job never ran. The common setup becomes `$tries = 0`, a `$maxExceptions` cap and `retryUntil()` as a backstop [7]. The author of PR #61737 described a worker caught in this loop, OOM-killed on Laravel Cloud [9].

The 13.34 counter is a small, well-made piece of engineering. The worker writes a marker to the cache when it picks up the job and deletes it when the attempt ends. If the marker is still there at the next pickup, the previous attempt died, and that death counts as one exception [11]. It costs two cache calls per attempt [12]. The marker is deleted only when an attempt ends. A deploy that outlasts its grace period and reaches SIGKILL should therefore leave one behind as well, so the same attribute ought to count deploy deaths [1]. The catch is storage. The cache has to be shared between workers. On the `array` driver, or `file` inside separate containers, the marker dies with the container [12].

The 13.33 flag needs more care. PR #61591 adds a static `Worker::$killOnTimeout`. Set it to false in `AppServiceProvider::boot()` and the worker survives the timeout, moving on without re-bootstrapping the framework or reopening every connection [13][14]. The discussion on PR #61622 raises two problems. An exception thrown from a signal handler propagates only when PHP returns to executing instructions, so a job stuck in a blocking system call never sees it [15]. If the exception does unwind, the worker carries on with whatever state the interrupted job left [16]. "SIGKILL is brutal, but it guarantees that state dies with the process," the post's author wrote [17]. A broad `try/catch (\Throwable)` in `handle()` will also swallow the exception and make the timeout disappear. The post gives these risks as the reason the default stays true [18].

Neither release changes `stop_grace_period` [2]. Ten seconds is enough for twenty of the post's half-second email jobs [3][21]. For a job waiting on a model, the author wrote, "ten seconds isn't much" [20].

I run model-calling jobs on `$tries = 0` with a shared cache. In that setup I would put the 13.34 attribute on every such job. I would leave `killOnTimeout` at its default until every `handle()` method has been checked for catch-all blocks and blocking calls. Neither choice sets the timeout values themselves. The author wrote that neither flag replaces "a strict ordering of timeouts, from the innermost to the outermost," and added: "In years of running Laravel queues this is the part I've seen get wrong most often" [19].

What to watch

  • A later 13.x release that makes crash counting a worker-wide default instead of an attribute every job class must carry.
  • Any change to the killOnTimeout default after the PR #61622 discussion about blocking system calls and leftover worker state.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories