Skip to content

Build1 publisher3 min readPublished

Reading the clock before SQLite's write lock made two AI-built queues issue expired leases

Both Astra + Luna runs of a SQLite job queue cleared their acceptance tests yet computed leases from a clock read taken before the write lock. An audit exposed it by holding one writer behind a barrier and advancing an injected clock during the other's wait.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Reading the clock before SQLite's write lock made two AI-built queues issue expired leases
Generated illustration

What happened

  • The test task was a durable TypeScript and SQLite reminder queue, built under four model configurations with two runs each, with Astra planning and reviewing in the paired setups.
  • With the clock moving from 0 to 10 during the lock wait and a 5-unit lease, both paired runs issued a lease ending at 5 instead of 15 and accepted expired ownership.
  • Astra + Luna had an estimated token cost 47.3% below Astra solo, and its runs took 46.3% longer in elapsed time.
  • The audit was retrospective, and three of its seven checks probe the same clock-after-lock defect.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint A green acceptance suite says little about lease correctness under contention, because sleep-based timing tests may never land on the lock boundary in CI.
  • exposure A lease that is expired when it is issued can be reclaimed at once under the queue's own rules, so a second worker can pick up a job the first worker is still running.
  • decision Adopting the cheaper paired setup means adding lock-wait tests with a barrier and an injected clock to acceptance, since the original checks let this defect through.
  • cost Anyone pricing the paired setup from this run gets a fixed-rate estimate built from token counts, which tells them nothing about an invoice or a subscription quota.

A lease is an owner token plus an expiration time [8]. A worker claims a job, does the work, then completes or fails it [8]. The database has to serialize those claims and transitions so two workers cannot act on the same live lease [8]. SQLite can make a writer wait for another transaction [9].

The post sketches the faulty ordering like this [9]:

```ts const now = clock.now(); await beginImmediateTransaction(db); // may wait for another writer await claimDueJob(db, now); await commit(db); ```

The timestamp is taken on the first line. If `beginImmediateTransaction` blocks, `now` ages for as long as the other writer holds the lock [9]. The claim is then evaluated against the stale value. Depending on the queue's rules, the result is a lease that is already expired, or a completion or failure accepted on ownership that lapsed during the wait [10].

The corrected ordering takes the transaction first and reads the clock once the write lock is held [11]:

```ts await beginImmediateTransaction(db); const now = clock.now(); await claimDueJob(db, now); await commit(db); ```

The diff moves one line down one slot. I'd expect a reviewer reading for logic to approve either version. The post attaches conditions to the fix: keep the time comparison and the state update in the same transaction, keep the owner-token checks, and account for clock semantics [11]. By the author's account, sampling after the lock removes this particular staleness and does not prove every timing or lease bug is solved [12].

A test has to force the wait to see any of this. Timing tests built on real sleeps can be flaky and may never reach the lock boundary in CI [16]. The post's ORCH-1 audit used two independent SQLite connections [13]. Connection A held an immediate write transaction behind a test barrier while connection B started a claim and reached the lock [13]. An injected clock then moved past the lease deadline before A was released, and the same shape ran against claim, complete and fail [13]. In the paired runs, the claim's lease end was computed from the pre-wait time, so all 10 units of the wait went missing [2]. The author recommends asserting state invariants per operation, not an incidental number of milliseconds [19].

"This is a narrow engineering result," the author wrote [18]. Only four of the seven audit checks look at anything other than clock ordering [4]. For the token saving to describe your bill, you would need to pay the same historical rates the estimate used and run a task that splits planning and implementation the way this one did [6][2]. The excerpt does not say whether the Astra solo runs read the clock in the same place.

"A useful review question is whether a time-sensitive decision happens before or after a transaction wait," the author wrote [17].

What to watch

  • Results for the Astra solo runs on the same three clock-after-lock checks; they would show whether the defect tracks the pairing or the task.
  • Publication of the ORCH-1 harness as reusable test code that other SQLite queue projects can run against their own claim, complete and fail paths.
  • A rerun with more than two runs per configuration, or on a second task, to test whether the token-cost gap between Astra solo and Astra + Luna holds.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories