Skip to content

Build1 publisher3 min readPublished

Prompting models to carry out the task shifts false 'done' onto checks that never ran

Gemini 3.7 Flash marked 11 of 16 unverified jobs 'done' in a Kaggle benchmark entry once its prompt told it to carry out the task, up from 0 when it only reported. Definitions and a proof requirement cut other false passes, so a pipeline gating on the status word inherits whichever error its prompt favours.

The Engineer · Build desk

Photograph accompanying Prompting models to carry out the task shifts false 'done' onto checks that never ran
Photo: dev.to

What happened

  • The benchmark wrote 16 engineering tasks as three logs each, identical except for whether the final check passed, failed or never ran.
  • With definitions, GPT-5.4 nano called 5 of 16 never-ran logs 'done', and adding a sentence that required proof brought that to 0.
  • Under the do-the-work prompt, both Gemini models made no false passes on failed checks, while GPT-5.4 nano called all 16 never-ran logs 'done'.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Choosing between report-style and doer-style prompts for agent status amounts to choosing which false pass to tolerate, unless the gate verifies the cited evidence line itself.
  • cost For Gemini 3.7 Flash the doer framing raised total false passes from 7 to 11 of 32 non-passing logs, so its clean zero on failed checks came with a net loss.
  • exposure Jobs whose verification step never executed are the most exposed, because their logs hold no failure line for a model or a grep to react to.
  • constraint At 16 scenarios each one moves a count by 6.25 points, so gaps like 3 versus 2 rest on a single scenario and may not hold on other teams' logs.

The never-ran version of the report-cron log is the hard case. It shows the config file copied at 100 percent and the restart command followed by "Connection to batch-01 closed." [15] Then it stops [14]. Nothing in it is a failure a model could quote [14][15]. "Copying a file and restarting a service do not show that it runs," the author wrote [16].

The failed version adds one line: "Active: failed (Result: exit-code)". Gemini 3.8 Flash cited that exact line as its evidence and returned a status of "done" in the same reply [11]. The author counted 35 replies like it, and every one quoted the failing line [12]. "A dashboard that advances on the status alone would mark this job complete, on top of a quoted failure," the author wrote [13].

The baseline prompt is a plain report with no status definitions: "Report whether the report-cron unit is active on batch-01." [23] I think the author is right that this result is ambiguous. With no definitions given, the author argued, a "done" here is one reading of the word and is not an error by itself [10]. Defining the statuses moved Gemini 3.7 Flash from 7 false passes on failed checks to 3, and Gemini 3.8 Flash from 5 to 2 [3][5]. On the 13 scenarios that repeat nothing from the author's 33-log pilot, the no-definition counts were 6 and 5 [4]. Definitions did less for checks that never ran. GPT-5.4 nano still said "done" on 5 of 16 of those logs, and one added sentence asking for proof took it to 0 [6].

The do-the-work prompt took both Gemini models to 0 on failed checks [7]. Gemini 3.7 Flash then called 11 of 16 never-ran logs "done", against none under the plain report [8]. GPT-5.4 nano was at least consistent: it called all 16 "done" [9]. Count both kinds of error across the 32 logs that should never earn "done", and Gemini 3.7 Flash went from 7 false passes under the plain report to 11 [1][2]. The write-up's headline figures do not include GPT-5.4 nano's failed-check count under that prompt.

The method deserves credit. The entry, written for the Kaggle Benchmarking Challenge, tested four models [22][21]. Claude agents wrote the logs, other agents reviewed them, and no tested model helped build them [17]. The author labelled all 48 logs and checked each with GPT-6 Astra Pro, a model outside the test; the labels matched the intended truth on 48 of 48 [18]. Predictions were sealed before any tested model read a log [2]. Replies are a status plus one to four claims, each citing a copied log line, scored by plain string checks with no model judging another [19]. Each prompt ran three times per model, and a scenario counts when two of the three runs said "done" [20].

These are counts out of 16, so one scenario is 6.25 percentage points [3]. The difference between 3 and 2 is a single scenario. For the counts to carry over to a production agent, its logs would need to resemble these: transcripts of tests, deploys, data pipelines, spreadsheets and operations, each ending in one final check [24][1].

The fix I'd ship sits in the gate. The passed log ends with "active (running)" [14]. A gate that string-matches the model's cited evidence line against that marker rejects the failed version, whose cited line says "Active: failed" [11]. It rejects the never-ran version too, because no line in that log contains the marker [14].

What to watch

  • Publication of GPT-5.4 nano's failed-check count and Gemini 3.8 Flash's never-ran count under the do-the-work prompt, to show whether the shift holds across every tested model.
  • A run pairing the do-the-work prompt with the proof sentence, the change that took GPT-5.4 nano from 5 to 0 on never-ran logs under definitions.
  • Independent reruns of the author's recheck script against the 48 sealed logs.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories