Skip to content

Build1 publisher3 min readPublished

41 queued tasks, zero ready to dispatch: the agent bottleneck was judgment

A one-week, two-repo field report from a solo developer: implementation capacity sat idle because nobody was re-judging stale tasks. Two of seven hand-dispatched tasks had already rotted.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying 41 queued tasks, zero ready to dispatch: the agent bottleneck was judgment
Generated illustration

What happened

  • The author's ledger held 41 tasks across two repositories, and never shrank despite daily use of AI agents.
  • Running the author's homegrown 'list the tasks that are ready to start' command produced no output: zero of the 41 tasks were ready to hand to an implementation session, leaving implementation capacity idle; the author concludes the bottleneck was judgment.
  • The report is a record of two repositories, one week, one person (n=1), with every number taken from logs and commits, and the author says to generalize only within that range.
  • The day before building the loop, the author hand-dispatched 7 tasks to implementation sessions as a trial.
  • Two of the 7 hand-dispatched tasks had premises that had already collapsed by the time work started: a spec they depended on had changed, or the problem itself had been dissolved by some other change.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A developer running AI coding agents across two repositories found 41 open tasks in his ledger, ran his own "list the tasks that are ready to start" command, and got no output at all [1][2]. Implementation capacity was idle not because the build sessions were slow, but because nothing in the queue was in a shape anyone could hand to one [2].

The evidence arrived a day before the loop existed. He hand-dispatched 7 tasks to implementation sessions as a trial, and 2 of them had premises that had already collapsed by the time work started: a spec they depended on had changed, or the problem had been dissolved by some other change [4][5]. That is 29 percent of a very small sample [6], and the mechanism is unremarkable. Ledger entries encode the world as it was on the day they were written, so anything that has sat for weeks needs a fresh "is this still worth doing" pass before implementation [24]. Skip that layer, the author argues, and the agent stacks correct code on rotten premises [25].

The rebuild split task processing into three roles, judge, build, and human [7], with deliberately thin plumbing. No pull requests: build sessions stack commits on a task branch, the judge session accepts by reading git diff --stat and re-running the tests, and the fast-forward merge happens only after the human says merge [8][9]. The list of unmerged branches is the acceptance queue [10].

The balance sheet for the first full day with all three roles running: open tasks fell from 41 to 22 across the two repositories, 13 to 5 and 28 to 17 [11]. 27 tasks closed against 13 merges, and the two are not 1-to-1 because some merges were docs-only and some tasks closed with no merge at all [12]. 6 new tasks were filed during the cycle, each approved by the human on the spot [13], and 2 of the 27 closures belonged to an adjacent repository outside the original 41 [14]. Net, 25 of the 41 went away, 61 percent [23].

The distribution matters more than the total. The judgment pass closed 6 tasks by decision or withdrawal before anything was dispatched, and read-only investigation sessions closed 4 more [15][16], so 10 of the 27 closures, 37 percent, needed no code written [17]. Judgment also repaired ledger state rather than closing it: three tasks were marked blocked although their dependency had completed weeks earlier, and one carried a resume condition that could never structurally fire [18]. Those moved forward on a state correction alone [18].

The build side was not the weak link. 14 implementation sessions ran that day under a parallelism cap of 3 and produced the 13 merges [19]. The missing capacity was the layer that turns a stale entry into a dispatchable one [2].

One detail worth copying: every kickoff packet he wrote on day one was wrong somewhere, either a premise that was only half true or a prescribed fix that would not actually close the hole [20]. The countermeasure was to put a Phase 0 instruction at the top of every packet, re-verify the premises and stop and report instead of implementing if they are falsified, which turns the packet from an order into a hypothesis [21].

Two caveats the author states himself. This is n=1, one person, one week, two repositories, with numbers taken from logs and commits [3]. And the day being counted was still operated manually; unattended mode came later [22]. So it is evidence about where a queue jams, not yet about a loop that ran without a human in it.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories