Build1 publisher3 min readPublished
The prompt never arrived: a Windows batch shim was worth 15 of 24 runs in an agent eval
A harness scored 3 of 24 until its operator stopped trusting shutil.which. The variable under test turned out to be the plumbing, not the model.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The account was published on dev.to under the headline "A Batch-File Shim Was Truncating My Agent's Prompts on Windows (3/24 to 21/24)", with the environment given as Windows 11 Home, Python 3.x, and Claude Code installed via npm install -g @anthropic-ai/claude-code, measured 2026-08-14 to 2026-08-16.
- Harness A used the computer use API directly with claude-sonnet-5 and max_turns=40, running 8 tasks x 1 trial.
- Harness B used Claude Code headless (claude -p) plus Playwright with the same model, running 8 tasks x 3 trials = 24 runs.
- Success in both harnesses was decided by machine verification only: submitted form payloads compared against expected JSONL, downloaded files checked for existence and size, answers matched by regex; the agent's own claim of "done" was never trusted as a success signal.
- The first run of harness B scored 3 out of 24, with only the first task passing.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A developer running a web-task eval on Windows 11 watched the same model, on the same prompt wording, move from 3 out of 24 to 21 out of 24 after changing only how the prompt reached the process [6]. The mechanism was a batch-file wrapper installed by npm that silently discarded everything after the first newline of a multi-line prompt [7][8].
The setup, per the dev.to writeup [0], was two harnesses over the same eight tasks: harness A calling the computer use API directly with claude-sonnet-5 and max_turns=40, one trial per task [2]; harness B running Claude Code headless via `claude -p` plus Playwright, same model, three trials per task for 24 runs [3]. Grading was machine-only: form payloads diffed against expected JSONL, downloaded files checked for existence and size, answers matched by regex, and the agent's own claim of "done" never counted as a pass [4]. That is the right way to build a scoreboard, and it still did not protect the result, because the failure was upstream of anything the grader could see.
Root cause: the npm global install drops a file called `claude.CMD` on PATH, and that is what `shutil.which("claude")` resolves to rather than the real binary. Its body is a one-line forward to `claude.exe` with `%*` [8]. The author is careful about the boundary between measurement and inference. Measured: swap only argv[0], hold prompt, flags and environment constant, and the tail either arrives or does not [9]. Inferred, and explicitly not isolated: `%*` is cmd.exe's argument forwarding and the newline appears to be lost somewhere in that path, though the author did not determine which layer drops it and stopped once either workaround sufficed [10]. The reproduction is four lines of subprocess and a marker string; through the shim, the model replies asking for the missing instruction, having seen only line one [11].
The headline number deserves a discount. The post's own progression shows 3/24 came with a prompt that never included the target URL at all, an error harness A masked because it navigated the browser programmatically before the agent's turn [14][15]. Adding the URL got it to 6/24, and at that point only tasks whose instruction fit on a single line passed [16]. So the transport bug is worth the gap from 6 to 21, or 15 of 24 runs, about 62.5 points of success rate [17]. That is still the largest single term in the result, and it is a term with no model content in it whatsoever.
What makes this a class of defect rather than one person's bad afternoon: there is no exception, no non-zero exit code, and no warning anywhere in the pipeline [13]. A truncated prompt is a well-formed prompt. The only signal is a low score, which is indistinguishable from "the model cannot do this," and the author admits to suspecting the model first [13]. Any published Windows agent benchmark that shells out to an npm-installed CLI and has not printed the bytes the child process actually received is, in my judgement, an unvalidated number. The check is cheap: send a two-line prompt with a marker on line two and assert the marker comes back.
Watch for the portability trap in the fix. Resolving `claude.exe` directly works; normalising newlines to spaces is the more portable option because it sidesteps the shim on any platform [12]. Harnesses that pick the first fix inherit a path assumption that the next installer layout will break.