Skip to content

Build1 publisherNot yet confirmed elsewhere3 min readPublished

132 blockers, three defect families: the bigger model wrote better prose and the same bad plans

A 157-goal field test of an LLM plan-and-critique loop found failures clustered in three structural families. Swapping in gpt-4o changed the writing, not the dependency graph.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying 132 blockers, three defect families: the bigger model wrote better prose and the same bad plans
Generated illustration

What happened

  • A field test of an LLM plan-and-critique engine over 157 goals logged 132 concrete blockers across the 63 goals judged under strict review.
  • The revision loop stalled at a median of two revisions across 33 strict goals, at which point the convergence detector fired and the engine escalated.
  • The author proposes a deterministic precondition linter and credits it with removing 64 of the 132 blockers.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint For this defect class the model upgrade path is closed, so the quality ceiling of a plan sits in the harness the builder writes rather than in the vendor's next checkpoint.
  • cost Each escalated goal is paid for twice over in plan generation and critique before the engine concedes, and the spend buys reshuffled task order rather than a closed graph.
  • decision Teams running plan-then-review loops now have a concrete alternative to another revision round: spend the engineering on a symbolic check of preconditions and keep the LLM for drafting.
  • contradiction The published counts cover 121 of 132 blockers, so the tidy three-family story needs a footnote before anyone treats it as a complete taxonomy of planner failure.

The taxonomy does not close. Unverified dependencies at 57 blockers, unsafe sequencing at 46, and weak rollback at 18 sum to 121 [6][7][8][14], against 132 blockers reported across 63 strict goals [5]. Eleven sit outside the three named buckets [14], which is awkward next to the claim that every strict goal that failed did so because of one of the three families [18]. Either those eleven are extra instances inside goals already accounted for, or there is a fourth family that has not yet earned a name.

The proposed fix has a similar seam. A deterministic precondition closer, run on the draft to verify that every declared precondition is produced by some earlier task, is credited with removing 64 of the 132 blockers [10][19]. But unverified dependencies only number 57 [6], so seven of that 64 have to be drawn from the sequencing pile [15]. That is coherent rather than sloppy: once you know which task produces `replica_verified`, you also know the backfill cannot legally precede the quality gate that establishes it [7]. Dependency closure and ordering are one check read in two directions. Rollback gets nothing from the pass [8], and 68 blockers survive it [20].

The reason parameters do not help is visible in the rollback example. In `db-10-multi-tenant-split`, the dual-write task did have a rollback; it switched back to single-write and left the inconsistencies that dual-write had introduced [13]. That is not a model that lacks the word "rollback." It is a model that generates each step against local plausibility while the property being violated is global to the plan. Author Debashish Ghosal reports testing gpt-4o as planner with a mini critic and then for both roles, and got the same defect pattern with better prose [1]. The revision loop behaves the way you would expect from a local generator: it fixes one blocker and introduces another, reshuffles order without closing the gap, adds rollback to the wrong task [9].

What makes the diagnosis load-bearing is the critic's stability. The same blockers were found again across revisions, so the loop's failure is attributed to the planner's inability to structurally repair rather than to reviewer drift [2]. Without that, planner incapacity and critic noise would be indistinguishable, and the convergence detector firing at a median of two revisions across 33 strict goals [9] would read as flakiness instead of a wall.

Scope, honestly stated: this is one engine [3], one author's labels, one goal set. Strict goals were 63 of 157 [17], and the failing ones averaged 2.1 blockers each [16]. The three families are a finding about that harness until someone reproduces them elsewhere. The claim that survives regardless is narrower and more useful than the headline: a plan is a graph with checkable properties, and checking is cheaper than hoping the next checkpoint has learned to plan.

What to watch

  • Whether a shipped precondition closer actually removes 64 blockers on a rerun of the same 63 strict goals, or a materially smaller number.
  • Whether the 11 blockers outside the three-family total are duplicates within counted goals or a fourth defect family.
  • Whether the same three families show up in a different planner on an independent goal set, which is what would turn one project's log into a taxonomy.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories