Skip to content

Build1 publisher3 min readPublished

A dead Redis lock made Munchable's scheduler log twelve skipped cron runs as successes

Munchable's scheduler logged twelve successful cron runs in three hours on 14 September 2026 while a stale Redis lock made each one skip. Each skip answered HTTP 200, the status the platform records as a successful run whether or not any curation happened.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying A dead Redis lock made Munchable's scheduler log twelve skipped cron runs as successes
Generated illustration

What happened

  • The overlap lock is a Redis key set with NX, and its expiry first matched the job's worst-case runtime: three hours on the schedule, six from the command line.
  • A run that is cancelled or killed never reaches its finally block, so the two runs stopped by hand left their lock key in place for the full three hours.
  • The replacement is a lease: the key now expires after 300 seconds unless the running job renews it, and the job renews it every 90 seconds.
  • Every scheduled route also fails closed on auth, answering 503 and refusing to run when its secret is missing or shorter than 16 characters.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure A cron route is a public URL, so a job that runs when its secret is unset can be set off by whoever finds the path, a crawler on a forgotten preview deploy included.
  • constraint Sizing a lock's expiry to the longest run means every killed run blocks the schedule for that full length, so the TTL is an availability setting as much as a safety one.
  • cost Under the lease, curation can occasionally run twice; a job whose duplicate run sends email or moves money would need the code to prefer the opposite failure.
  • decision Platform run history tracks HTTP status, so a deliberate skip and a finished job look identical there, and proof that work happened has to come from the job's output.

"A cron job on a serverless platform is an HTTP endpoint with a timer pointed at it," Munchable's post on dev.to says [3]. Anyone who guesses the path can call it, and it can be invoked twice at once [3]. Without a token, the company's indexnow route answers `{"error":"unauthorized"}`, a 401 [1].

The auth check returns one of three states, and the third is the one I would copy [4]. It covers a secret that was never configured [5]. Running open in that case is, according to the post, "how a job that spends money ends up invocable by a search crawler in a preview environment nobody remembered deploying" [6]. The token comparison is constant time [7]. A length check runs first because `timingSafeEqual` throws on inputs of different lengths, and the post concedes this leaks the secret's length [7]. With a 16-character floor on the secret, I would accept that leak too [5].

The lock failure came from a rule that was right as far as it went. The original expiry followed from the premise that "the lock must outlive the run, so the TTL has to be at least as long as the longest run" [15]. The rule did not account for a run that dies before cleanup [9]. On the 15-minute schedule the runner used at the time, a three-hour key spans twelve ticks [1]. Each tick got back `{"skipped": true, "reason": "locked"}` [10].

The lease ties the lock to liveness. A live job keeps the lock for as long as it keeps renewing, and a dead one loses it within five minutes [13]. The post says three renewals fit in each lease, so one failed renewal is a non-event [12]. The margin is wider than the post claims. After any successful renewal, the next attempts land at 90, 180 and 270 seconds, all inside the 300-second expiry, so a live job loses the lock only after three consecutive failures [3]. The longest a dead job can block the schedule falls from 180 minutes to five, a factor of 36 [2]. The curation runner also shares the key with its command-line equivalent, so an operator running it by hand cannot collide with the schedule [11].

Two smaller details hold up under review. The renewal timer is unref'd, so a pending renewal cannot keep a finished worker alive, and on a platform that bills by duration a process that will not exit costs money [16]. The most useful line in the change, which is not something I often say about a catch block, is the comment on the one that swallows renewal errors. It argues that losing the lock entirely risks only an overlap, and that an overlap is the lesser failure next to curation stopping for hours [14].

The fix as described changes how long a dead lock lives [13]. The post does not describe a change to the 200 that a locked run returns [10]. If that response is unchanged, a dead lease can still produce at most one green tick that did no work on a 15-minute schedule, because its five-minute window is shorter than the interval [4].

What to watch

  • Whether Munchable changes the locked-skip response from HTTP 200 to a status or alert the scheduler treats as a failure.
  • Whether a Redis outage long enough to miss three renewals produces the overlapping curation run the code comment accepts.
  • Whether Munchable publishes how it now checks curation output independently of the scheduler's run status.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories