Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished

A 170-goal agent field test costs $0.49. Proving it actually passed costs more.

PlannerCritic's sweeps are cheap enough to gate every release. The harder problem is telling a clean run from a rig that quietly counted nothing.

The Engineer · Build desk

How we use AISend a correction

Photograph accompanying A 170-goal agent field test costs $0.49. Proving it actually passed costs more.
Photo: metr.org

What happened

  • The first PlannerCritic field test ran 157 goals across 35 domains for $0.30 in an hour and returned 10 issues, only one of them a plain failure.
  • v0.2.1 ran 170 goals for $0.49, fixed 10 more code-review bugs, found no field-test issues, and added a live-critic boundary evaluator.
  • The author reports that code review, costed at $0 and two hours, found 31 bugs before any tokens were spent on the v0.2.0 sweep.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • capability At under three-tenths of a cent per goal, the sweep drops below the level where anyone has to approve it, which is what lets it sit in front of every release rather than every quarter.
  • constraint A reviewer that returns different blockers each revision means a green run cannot be replayed as proof; the gate samples the system instead of certifying it.
  • decision Ordering review before tokens moves the load onto human reading, so the release now depends on review quality holding at tens of bugs per cycle rather than on the sweep catching anything.
  • contradiction The published goal additions do not sum to the published total, so the 10-to-0 improvement rests on two suites that may not contain the same work.

Zero failures is a claim about the harness before it is a claim about the agent, and the v0.1.0 run of PlannerCritic [6] had four separate ways to return a number that meant nothing. Eighty-eight percent of the assertion files were in a format the runner did not recognise, and it returned 0/0 rather than an error [23][1]. The budget and replan dimensions also reported 0/0, because dispatch passed five arguments to a function that takes four [9]. The in-memory SQLite store reset between runs, losing cross-dimension state [11]. The results parser read `trace.get("status")` instead of `trace["result"]["status"]`, converting five passes into reported failures [10]. On the author's own breakdown, four of the ten findings were in the measuring apparatus rather than the planner [25]; exactly one was the sort of defect a conventional test suite would have called a failure [8].

That is the reason a green sweep needs something other than a pass rate to be readable. Three of those four bugs produced zeros or inversions that no summary line would flag. What distinguishes a clean pass from a silent one is the denominator: assertions evaluated per dimension, published next to the result, plus at least one goal that is supposed to fail and whose failure the rig has to report. Neither release note describes such a canary [3][17].

The economics are more interesting once you divide. v0.1.0 spent about $0.0019 per goal [19]; v0.2.1 spends about $0.0029 [20]. Per-goal cost rose roughly half while the suite grew about eight percent [21], and the source does not say why, though v0.2.1 added a live-critic boundary evaluator and an operational benchmark [17]. More usefully: the first run cost about three cents per issue found [22], and that metric goes undefined the moment the gate starts returning zero. What replaces it in the source is the assertion that a sweep is cheaper than debugging one production incident caused by a bad plan [27], which is not a number you can check.

Two findings put a floor under the price and a ceiling over the fix. Local Qwen3.5-4B and 9B models could not produce structured JSON [14], so the run cannot be moved off a paid API to get to zero marginal cost. And GPT-4o reproduced the same defect patterns as gpt-4o-mini [16], so the planner gap is prompt and schema work, not a billing decision. Meanwhile the critic is non-deterministic: strict goals re-run at `revision_cap=4` all escalated, with different blockers on each revision [15].

The accounting deserves one more look. v0.2.0 is described as 170 goals across 40 domains against v0.1.0's 157 across 35 [3][7], with 14 new goals in five new enterprise domains and three new adversarial-policy goals [18]. Seventeen additions on a base of 157 gives 174, not 170, leaving four goals unaccounted for [26]. The whole 10-to-0 arc [5] assumes the two suites are comparable, and a four-goal discrepancy in the published totals is the same class of problem as the assertion files: the number is stated, and nothing in the write-up demonstrates it was counted.

What to watch

  • Whether the project publishes per-dimension assertion counts next to pass rates, which is the minimum needed to audit a zero.
  • Whether a deliberately failing canary goal enters the suite so the harness has to demonstrate it can still report a failure.
  • Whether the four-goal gap between stated additions and the stated 170 total is reconciled in later v0.2.x notes.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence34
Adoption11
Hype gap+24
Incentives68
Confidence33
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    57 of 65 assertion files were in the wrong format and the harness did not notice, quietly returning 0/0 results with no crash and no error.

  2. [2]

    In v0.1.0, a result of 0 failures meant the harness was silently broken; the author writes that a field test returning 0 failures should first be distrusted.

  3. [3]

    v0.2.0 ran 170 goals across 40 domains, fixed 31 bugs found in code review, and the field test itself found zero new issues.

Sources

1 independent publisher whose own reporting we read for this story.

  1. dev.to

    1 article · August 24, 2026

    I Ran 170 Agent Goals for $0.49. The Field Test Found 10 Issues That Unit Tests Never Would.

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Loading related stories