Skip to content

Build1 publisher3 min readPublished

launchd Tells You Nothing When a Job Dies, So Your Revenue Reports It Instead

A dev.to writeup documents a six-day silent crash loop across 26 scheduled macOS jobs, and a single-command health check built around one useful contract: exit 1 if anything is red.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • A dev.to post headlined "Six Days of a Silent Crash Loop: One Command That Health-Checks 26 launchd Jobs" describes the author discovering, on the afternoon of a layoff, that half the automation supporting his side income had quietly stopped running and nobody, including him, had noticed.
  • macOS launchd, the daemon-management layer, responds to a crashing script by saying nothing and waiting for the next scheduled run.
  • Even if a job is returning exit 78, nobody finds out unless you run launchctl list yourself.
  • In the author's environment, the job com.shun.agentmemory sat in a crash loop for six days in June 2026.
  • A worse half-alive state occurred: the port was open but no worker was behind it, so every HTTP request returned 404 while the process still existed.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

An engineer who had just been laid off spent the same afternoon finding that half the automation propping up his side income had quietly stopped running, and that macOS had said nothing about any of it [1][9]. The mechanism deserves attention from anyone running scheduled work: launchd answers a crashing script with silence and waits for the next scheduled run [2].

According to the post, a job can return exit 78 and nobody learns about it unless a human runs `launchctl list` [3]. In the author's own environment, `com.shun.agentmemory` sat in a crash loop for six days in June 2026 [4]. The worse state was the half-alive one: the port was open, no worker was behind it, and every HTTP request returned 404 while the process continued to exist [5]. A liveness check that only confirms the process exists misses that entirely [6].

The scaling argument is the part that generalises. Three launchd jobs get eyeballed every morning; ten get checked weekly; at 26 you stop checking altogether [8]. The author puts the threshold at roughly 20 jobs, past which silent death becomes routine [7]. That is not a discipline problem, it is an attention budget, and the discovery channel in this case was revenue reaching zero [10].

The response is a script, `automation-health.sh`, with a contract worth copying even if you never see the code: exit 0 means all green or warnings only, exit 1 means at least one red [16]. That single integer is what lets the check run without a human deciding to run it, wired into a StopHook that fires when a Claude session ends, or into cron each morning [17]. The script covers nine sections by the author's count, each reporting green, yellow, or red [15], though the accompanying diagram actually enumerates ten blocks because of a 5.5 [28].

Two of those checks are worth stealing. First, the launchd section loops every plist matching `com.shun.*` or `com.lily.*`, reconciles against `launchctl list`, re-injects anything unloaded with `launchctl bootstrap`, and goes red both when re-injection fails and when the previous exit code was nonzero [18]. That pair is what a process check cannot see. Second, the agentmemory section requires launchctl status and an HTTP 200 from `http://localhost:3111/agentmemory/health`, and goes red if either is missing [23]. That is the direct answer to the open-port-no-worker case.

The rest is estate hygiene: hook scripts missing versus merely non-executable [19], log freshness windows of 48 hours for skill-harvest and 24 hours for conversation logs [20][21], Obsidian markers and index coverage [22], memory-layer files plus duplicate-burst detection [24], a 5GB ceiling on `~/.claude` [25], seven weekly and monthly batches checked against time budgets [26], and a check that no script is registered in both cron and launchd, which is the double-execution bug you inherit after a migration [27].

Caveats: this is one person's machine, self-reported. The income figures, ¥600,000 per month before and ¥1.2M per month after, are unverifiable and roughly a doubling [9][11][29]. What is portable is narrower and more useful: an exit code, a scheduler that runs the check for you, and assertions on responses rather than on the existence of a process.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories