Build1 publisher3 min readPublished
DeepSeek and Qwen3.8-Max score near a blind guesser on macOS bash 3.2's set -u cases
DeepSeek and Qwen3.8-Max scored 0.52 and 0.40 on bash 3.2's set -u cases in a 57-script macOS benchmark where a script-blind stub scores 0.446. The ground truth came from running macOS's own /bin/bash, the same check that exposed a scoring flaw in the author's harness.
The Engineer · Build desk
What happened
- The author ran 57 small scripts once each through /bin/bash 3.2.57 on macOS 27.0.1 under env -i, keeping exit status and last stdout line as ground truth.
- Of the 57 scripts, 31 exit 0 and 17 of those end by printing the word end, so a stub that never reads the script collects a lot of credit.
- Kimi, DeepSeek and Qwen3.8-Max stopped after the first 19-case chunk when the included usage quota ran out; only Qwen3.8-Flash ran all 57.
- DeepSeek and Qwen3.8-Max scored a perfect 1.00 on every behaviour family except set -u.
- Six expected answers began with a scripts/ path the prompt never showed the models, capping scores on those cases until a script_path field was added.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- contradiction By the post's own floor, Qwen3.8-Max did worse than a script-blind stub on set -u, so the 'barely above' description fits DeepSeek alone.
- exposure A script written against bash 4.4 or later that expands an empty array under set -u aborts on a stock Mac's /bin/bash, and the tested models were least reliable on that construct.
- constraint Any benchmark that grades bash 3.2 error text has to show the model the invocation path, or it holds every model below a perfect score.
The 0.446 floor can be rebuilt from the counts the post publishes. A stub that always answers exit 0 and "end" gets the full 1.0 on 17 cases and 0.6 on the other 14 zero-exit cases [6] [8]. It earns nothing on the 26 scripts that exit non-zero [4]. The sum is (17 + 8.4) / 57, or 0.446 [1]. A model at 0.89 is not 89% of the way to understanding bash, the post says: "it is 0.44 above a constant-string stub." [9] Publishing that floor before any model result is good practice. So is the plumbing. The baseline script imports the scorer from the notebook source so the two cannot drift [7], and the labels are parsed from captured interpreter output with nothing entered by hand [5].
The post lists three places where the shipped shell differs from the bash people learn on Linux: set -e being ignored inside a function whose call is being tested, local x=$(false) reporting success with x left empty, and set -u aborting on an empty array [2]. Only the last comes with a version number. Bash 4.4 made "${arr[@]}" on an existing empty array legal under set -u [3]. That family is where DeepSeek and Qwen3.8-Max lost their points [12]. The post's diagnosis is blunt: "That is a training-data artifact with a version number on it." [13] It is a reasonable guess. The harness scores predicted exit codes and last lines [6]. It has no view of what any model was trained on.
Against the floor, the two family scores split. Qwen3.8-Max's 0.40 is 0.046 below the whole-dataset 0.446, and DeepSeek's 0.52 is 0.074 above it [3]. The post describes both as "barely above the 0.446 floor" [23]. It does not say how many of the 19 shared cases are set -u cases, or what a stub scores on that family alone.
The harness bug came from the interpreter's own output format. Bash 3.2 prefixes its error messages with the path the script was invoked with [15]. In case e67 the measured stderr was "scripts/e67_setu_undefined_scalar.sh: line 6: undefined_scalar: unbound variable" [14]. All four models predicted exit 1, and none produced that line [14]. The prompt had said only "Invocation: ... /bin/bash <file>" [15]. The affected cases were grading a guess at the author's directory layout [15]. If no model could make that guess, the best possible score on the full set was 1 - (6 x 0.4) / 57, about 0.958 [2]. The fix put script_path in the prompt and left the measured labels untouched [16].
The method transfers better than the scores. The ground truth is whatever env -i PATH=/usr/bin:/bin /bin/bash printed on macOS 27.0.1 [4]. The author's prompt was wrong about what bash 3.2 would print, and the captured output is what showed it [15]. For the model numbers to carry over to reviewing real scripts, those scripts would need to look like these 57: small cases that each isolate one difference, read closed-book [4] [10]. The claim that the differences "bite production scripts" is the post's own [18]. According to the post, an autonomous agent working for MonkeyRun ran every command [17], and the scripts, measurements and harness are published for anyone to rerun [20].
What to watch
- Full 57-case runs for Kimi, DeepSeek and Qwen3.8-Max, the only basis on which the author says a ranking would hold.
- Rescored results after the script_path fix, showing whether the six path-prefixed cases now reach full marks.
- A stub floor computed for the set -u family alone, to settle whether Qwen3.8-Max's 0.40 is worse than a stub on those cases.