Build1 publisher3 min readPublished
A small Node verifier rejects an agent's saved report that totals 13 units instead of 16
A dev.to tutorial's Node verifier rejects an agent-written inventory report that claims 13 units when its rows, 7 and 9, add up to 16. Its expected values come from the trusted task input, so the test for done exists before the agent can announce it.
The Engineer · Build desk

What happened
- In the tutorial's example the file write genuinely succeeds, and a second successful write of the same report would leave the total just as wrong.
- The verifier checks report ID, revision, row shape, duplicates, rows against expected values and the computed total, returning a named reason code for each failure.
- It records a SHA-256 hash of the checked bytes so a later step can identify the same artifact; the author says the hash does not make the bytes correct.
- For handoffs, the tutorial swaps a prose note for a checkpoint marked completed: false that names the last verified problem and the next action.
- The whole demo runs under Node with no API key, paid service or model involved.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Agent runners that mark a task done on tool-call status need a separate completion check that only the verifier can pass, defined from the task before the agent runs.
- constraint The acceptance check is only as strict as the trusted input: where the task does not fix the expected rows, nothing stops a report that agrees with itself but not with the request.
- cost Adopting the checkpoint means keeping the expected input, artifact reference and verification receipt in private state for every unfinished task, alongside the public checkpoint.
Calling verifyReport(bytes, expected) runs its checks in a fixed order [7]:
1. Parse the bytes as JSON or return invalid_json. Anything that is not a plain object returns invalid_report. 2. Compare report_id and revision with the expected values. A mismatch adds wrong_scope. 3. Check the rows: an array of at most 100 entries, each with a string item and a non-negative safe integer for units. Anything else returns invalid_rows. 4. Flag repeated item names as duplicate_item. Flag any row set that differs from the expected rows as wrong_rows. 5. Sum the units and compare the sum with total_units. A mismatch adds wrong_total.
Only the parse and shape failures return early. The rest collect in a reasons array, so one call can report wrong_scope and wrong_total together [8]. Against the delivered file, the demo gets exactly one reason, wrong_total, and a total of 16 passes [10]. The delivered report is 3 units short [17].
Step 5 on its own is a self-consistency test. It sums the report's own rows. A report listing washers at 6 with a total of 13 would pass it, and only the step 4 comparison against the task's rows would reject it [19]. The trusted input is what turns a sanity check into an acceptance check. "The expected rows come from the trusted task input, not from the generated file asking to be trusted," the author wrote [6].
The check also has to run on the right bytes. The demo builds its buffer in memory. The tutorial says a real workflow should read the delivered path with readFileSync, and warns against verifying one in-memory object while assuming a different saved file contains it [11]. The receipt's hash [9] identifies the artifact on disk only if the bytes hashed came from disk.
The restart case is where I think the tutorial is strongest. The handoff it calls dangerous is a paragraph that says "The export worked; finish up" [12]. Every word of that note is true. The replacement checkpoint records last_verified_problem as wrong_total [13], the same reason string the verifier returned [18]. The failure code moves from the function into the checkpoint without being paraphrased on the way. "The next process must re-read the file; the checkpoint describes what was observed before interruption, not what must still be true now," the author wrote [14].
For the pattern to work in production, two things have to hold. The expected output has to be computable from input the agent cannot edit, and the verifier has to read the delivered bytes. An inventory total meets both. The author is plain about scope: "It is an original teaching example, not an agent runtime or a security boundary for arbitrary uploads" [16].
What to watch
- A version of the verifier for outputs the task input cannot fully determine, such as prose summaries, would show whether the trusted-input rule extends past sums.
- Agent runtimes exposing a completed flag that only a verifier result can set, instead of inferring completion from tool-call status.