BuildNot yet confirmed elsewhere1 publisher3 min readPublished
A 170-goal agent field test costs $0.49. Proving it actually passed costs more.
PlannerCritic's sweeps are cheap enough to gate every release. The harder problem is telling a clean run from a rig that quietly counted nothing.
The Engineer · Build desk

What happened
- The first PlannerCritic field test ran 157 goals across 35 domains for $0.30 in an hour and returned 10 issues, only one of them a plain failure.
- v0.2.1 ran 170 goals for $0.49, fixed 10 more code-review bugs, found no field-test issues, and added a live-critic boundary evaluator.
- The author reports that code review, costed at $0 and two hours, found 31 bugs before any tokens were spent on the v0.2.0 sweep.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- capability At under three-tenths of a cent per goal, the sweep drops below the level where anyone has to approve it, which is what lets it sit in front of every release rather than every quarter.
- constraint A reviewer that returns different blockers each revision means a green run cannot be replayed as proof; the gate samples the system instead of certifying it.
- decision Ordering review before tokens moves the load onto human reading, so the release now depends on review quality holding at tens of bugs per cycle rather than on the sweep catching anything.
- contradiction The published goal additions do not sum to the published total, so the 10-to-0 improvement rests on two suites that may not contain the same work.
Zero failures is a claim about the harness before it is a claim about the agent, and the v0.1.0 run of PlannerCritic [6] had four separate ways to return a number that meant nothing. Eighty-eight percent of the assertion files were in a format the runner did not recognise, and it returned 0/0 rather than an error [23][1]. The budget and replan dimensions also reported 0/0, because dispatch passed five arguments to a function that takes four [9]. The in-memory SQLite store reset between runs, losing cross-dimension state [11]. The results parser read `trace.get("status")` instead of `trace["result"]["status"]`, converting five passes into reported failures [10]. On the author's own breakdown, four of the ten findings were in the measuring apparatus rather than the planner [25]; exactly one was the sort of defect a conventional test suite would have called a failure [8].
That is the reason a green sweep needs something other than a pass rate to be readable. Three of those four bugs produced zeros or inversions that no summary line would flag. What distinguishes a clean pass from a silent one is the denominator: assertions evaluated per dimension, published next to the result, plus at least one goal that is supposed to fail and whose failure the rig has to report. Neither release note describes such a canary [3][17].
The economics are more interesting once you divide. v0.1.0 spent about $0.0019 per goal [19]; v0.2.1 spends about $0.0029 [20]. Per-goal cost rose roughly half while the suite grew about eight percent [21], and the source does not say why, though v0.2.1 added a live-critic boundary evaluator and an operational benchmark [17]. More usefully: the first run cost about three cents per issue found [22], and that metric goes undefined the moment the gate starts returning zero. What replaces it in the source is the assertion that a sweep is cheaper than debugging one production incident caused by a bad plan [27], which is not a number you can check.
Two findings put a floor under the price and a ceiling over the fix. Local Qwen3.5-4B and 9B models could not produce structured JSON [14], so the run cannot be moved off a paid API to get to zero marginal cost. And GPT-4o reproduced the same defect patterns as gpt-4o-mini [16], so the planner gap is prompt and schema work, not a billing decision. Meanwhile the critic is non-deterministic: strict goals re-run at `revision_cap=4` all escalated, with different blockers on each revision [15].
The accounting deserves one more look. v0.2.0 is described as 170 goals across 40 domains against v0.1.0's 157 across 35 [3][7], with 14 new goals in five new enterprise domains and three new adversarial-policy goals [18]. Seventeen additions on a base of 157 gives 174, not 170, leaving four goals unaccounted for [26]. The whole 10-to-0 arc [5] assumes the two suites are comparable, and a four-goal discrepancy in the published totals is the same class of problem as the assertion files: the number is stated, and nothing in the write-up demonstrates it was counted.
What to watch
- Whether the project publishes per-dimension assertion counts next to pass rates, which is the minimum needed to audit a zero.
- Whether a deliberately failing canary goal enters the suite so the harness has to demonstrate it can still report a failure.
- Whether the four-goal gap between stated additions and the stated 170 total is reconciled in later v0.2.x notes.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence34
- Adoption11
- Hype gap+24
- Incentives68
- Confidence33
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
57 of 65 assertion files were in the wrong format and the harness did not notice, quietly returning 0/0 results with no crash and no error.
- [2]
In v0.1.0, a result of 0 failures meant the harness was silently broken; the author writes that a field test returning 0 failures should first be distrusted.
- [3]
v0.2.0 ran 170 goals across 40 domains, fixed 31 bugs found in code review, and the field test itself found zero new issues.
- [4]
The author states that code review before the field test cost $0 and 2 hours and found 31 bugs, meaning the field test validated fixes rather than discovering them.
- [5]
The author states that code review caught all 41 bugs before the LLM ran, turning the field test from a diagnostic sweep into an immutable regression gate.
- [6]
PlannerCritic is an open-source engine in which one LLM writes a plan and a second LLM reviews it.
- [7]
The v0.1.0 field test ran 157 goals across 35 domains for $0.30 in 60 minutes of wall-clock time and found 10 issues.
- [8]
Only one of the 10 v0.1.0 findings was a traditional failure: the planner prompt did not explain the branches schema, and the LLM responded with kind: "rollback" and arrays of task objects where strings were required.
- [9]
The dimension dispatch had a signature mismatch: run_budget() takes 4 arguments but dispatch passed 5, so the budget and replan dimensions showed 0/0.
- [10]
The results parser read trace.get("status") instead of trace["result"]["status"], and five PASS scenarios were misreported as FAIL.
- [11]
Cross-dimension state was lost because the in-memory SQLite store reset between runs.
- [12]
The preconditions gate was too strict: established_by expected a task ID or an env: prefix while the LLM wrote bare fact names, and unit tests were green.
- [13]
The critic severity contract was wrong: the critic blocked plans for completeness concerns instead of concrete defects.
- [14]
Local models Qwen3.5-4B and Qwen3.5-9B could not produce structured JSON.
- [15]
The LLM critic is non-deterministic: strict goals re-run with revision_cap=4 all escalated, and the critic found different blockers on each revision.
- [16]
A stronger model did not close the planner gap: GPT-4o produced the same defect patterns as gpt-4o-mini.
- [17]
v0.2.1 ran 170 goals for $0.49, fixed 10 more code-review bugs, found 0 field-test issues, and added a live-critic boundary evaluator and an operational benchmark.
- [18]
v0.2.0 scope grew about 8% with 14 new goals across five new enterprise domains (identity management, multi-agent operations, site reliability engineering, supply chain policy, and FinOps/greenfield) plus 3 new adversarial-policy goals.
- [19]
The v0.1.0 field test cost about $0.0019 per goal.
- [20]
The v0.2.1 field test cost about $0.0029 per goal.
- [21]
Per-goal cost rose about 51% between v0.1.0 and v0.2.1 while the goal count rose about 8%.
- [22]
The v0.1.0 run cost about three cents per issue found.
- [23]
88% of the v0.1.0 assertion files were in the wrong format.
- [24]
The 41 bugs caught by code review are the sum of 31 fixed in v0.2.0 and 10 fixed in v0.2.1.
- [25]
Four of the ten v0.1.0 findings sat in the test apparatus rather than the agent: the assertion-file format, the dispatch signature mismatch, the SQLite reset, and the parser JSON path.
- [26]
The stated v0.2.0 additions total 17 goals against a net increase of 13, leaving four goals unaccounted for.
- [27]
The author argues the field test is cheaper than debugging one production incident caused by a bad plan.
ReportedInsufficientSource: the article's author2 sources— create a free account to open themView cited source
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toI Ran 170 Agent Goals for $0.49. The Field Test Found 10 Issues That Unit Tests Never Would.
1 article · August 24, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.