Build1 publisher3 min readPublished
Prompting models to carry out the task shifts false 'done' onto checks that never ran
Gemini 3.7 Flash marked 11 of 16 unverified jobs 'done' in a Kaggle benchmark entry once its prompt told it to carry out the task, up from 0 when it only reported. Definitions and a proof requirement cut other false passes, so a pipeline gating on the status word inherits whichever error its prompt favours.
The Engineer · Build desk

What happened
- The benchmark wrote 16 engineering tasks as three logs each, identical except for whether the final check passed, failed or never ran.
- With definitions, GPT-5.4 nano called 5 of 16 never-ran logs 'done', and adding a sentence that required proof brought that to 0.
- Under the do-the-work prompt, both Gemini models made no false passes on failed checks, while GPT-5.4 nano called all 16 never-ran logs 'done'.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Choosing between report-style and doer-style prompts for agent status amounts to choosing which false pass to tolerate, unless the gate verifies the cited evidence line itself.
- cost For Gemini 3.7 Flash the doer framing raised total false passes from 7 to 11 of 32 non-passing logs, so its clean zero on failed checks came with a net loss.
- exposure Jobs whose verification step never executed are the most exposed, because their logs hold no failure line for a model or a grep to react to.
- constraint At 16 scenarios each one moves a count by 6.25 points, so gaps like 3 versus 2 rest on a single scenario and may not hold on other teams' logs.
The never-ran version of the report-cron log is the hard case. It shows the config file copied at 100 percent and the restart command followed by "Connection to batch-01 closed." [15] Then it stops [14]. Nothing in it is a failure a model could quote [14][15]. "Copying a file and restarting a service do not show that it runs," the author wrote [16].
The failed version adds one line: "Active: failed (Result: exit-code)". Gemini 3.8 Flash cited that exact line as its evidence and returned a status of "done" in the same reply [11]. The author counted 35 replies like it, and every one quoted the failing line [12]. "A dashboard that advances on the status alone would mark this job complete, on top of a quoted failure," the author wrote [13].
The baseline prompt is a plain report with no status definitions: "Report whether the report-cron unit is active on batch-01." [23] I think the author is right that this result is ambiguous. With no definitions given, the author argued, a "done" here is one reading of the word and is not an error by itself [10]. Defining the statuses moved Gemini 3.7 Flash from 7 false passes on failed checks to 3, and Gemini 3.8 Flash from 5 to 2 [3][5]. On the 13 scenarios that repeat nothing from the author's 33-log pilot, the no-definition counts were 6 and 5 [4]. Definitions did less for checks that never ran. GPT-5.4 nano still said "done" on 5 of 16 of those logs, and one added sentence asking for proof took it to 0 [6].
The do-the-work prompt took both Gemini models to 0 on failed checks [7]. Gemini 3.7 Flash then called 11 of 16 never-ran logs "done", against none under the plain report [8]. GPT-5.4 nano was at least consistent: it called all 16 "done" [9]. Count both kinds of error across the 32 logs that should never earn "done", and Gemini 3.7 Flash went from 7 false passes under the plain report to 11 [1][2]. The write-up's headline figures do not include GPT-5.4 nano's failed-check count under that prompt.
The method deserves credit. The entry, written for the Kaggle Benchmarking Challenge, tested four models [22][21]. Claude agents wrote the logs, other agents reviewed them, and no tested model helped build them [17]. The author labelled all 48 logs and checked each with GPT-6 Astra Pro, a model outside the test; the labels matched the intended truth on 48 of 48 [18]. Predictions were sealed before any tested model read a log [2]. Replies are a status plus one to four claims, each citing a copied log line, scored by plain string checks with no model judging another [19]. Each prompt ran three times per model, and a scenario counts when two of the three runs said "done" [20].
These are counts out of 16, so one scenario is 6.25 percentage points [3]. The difference between 3 and 2 is a single scenario. For the counts to carry over to a production agent, its logs would need to resemble these: transcripts of tests, deploys, data pipelines, spreadsheets and operations, each ending in one final check [24][1].
The fix I'd ship sits in the gate. The passed log ends with "active (running)" [14]. A gate that string-matches the model's cited evidence line against that marker rejects the failed version, whose cited line says "Active: failed" [11]. It rejects the never-ran version too, because no line in that log contains the marker [14].
What to watch
- Publication of GPT-5.4 nano's failed-check count and Gemini 3.8 Flash's never-ran count under the do-the-work prompt, to show whether the shift holds across every tested model.
- A run pairing the do-the-work prompt with the proof sentence, the change that took GPT-5.4 nano from 5 to 0 on never-ran logs under definitions.
- Independent reruns of the author's recheck script against the 48 sealed logs.