Skip to content

Build1 publisher3 min readPublished

DeepSeek and Qwen3.8-Max score near a blind guesser on macOS bash 3.2's set -u cases

DeepSeek and Qwen3.8-Max scored 0.52 and 0.40 on bash 3.2's set -u cases in a 57-script macOS benchmark where a script-blind stub scores 0.446. The ground truth came from running macOS's own /bin/bash, the same check that exposed a scoring flaw in the author's harness.

The Engineer · Build desk

What happened

  • The author ran 57 small scripts once each through /bin/bash 3.2.57 on macOS 27.0.1 under env -i, keeping exit status and last stdout line as ground truth.
  • Of the 57 scripts, 31 exit 0 and 17 of those end by printing the word end, so a stub that never reads the script collects a lot of credit.
  • Kimi, DeepSeek and Qwen3.8-Max stopped after the first 19-case chunk when the included usage quota ran out; only Qwen3.8-Flash ran all 57.
  • DeepSeek and Qwen3.8-Max scored a perfect 1.00 on every behaviour family except set -u.
  • Six expected answers began with a scripts/ path the prompt never showed the models, capping scores on those cases until a script_path field was added.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • contradiction By the post's own floor, Qwen3.8-Max did worse than a script-blind stub on set -u, so the 'barely above' description fits DeepSeek alone.
  • exposure A script written against bash 4.4 or later that expands an empty array under set -u aborts on a stock Mac's /bin/bash, and the tested models were least reliable on that construct.
  • constraint Any benchmark that grades bash 3.2 error text has to show the model the invocation path, or it holds every model below a perfect score.

The 0.446 floor can be rebuilt from the counts the post publishes. A stub that always answers exit 0 and "end" gets the full 1.0 on 17 cases and 0.6 on the other 14 zero-exit cases [6] [8]. It earns nothing on the 26 scripts that exit non-zero [4]. The sum is (17 + 8.4) / 57, or 0.446 [1]. A model at 0.89 is not 89% of the way to understanding bash, the post says: "it is 0.44 above a constant-string stub." [9] Publishing that floor before any model result is good practice. So is the plumbing. The baseline script imports the scorer from the notebook source so the two cannot drift [7], and the labels are parsed from captured interpreter output with nothing entered by hand [5].

The post lists three places where the shipped shell differs from the bash people learn on Linux: set -e being ignored inside a function whose call is being tested, local x=$(false) reporting success with x left empty, and set -u aborting on an empty array [2]. Only the last comes with a version number. Bash 4.4 made "${arr[@]}" on an existing empty array legal under set -u [3]. That family is where DeepSeek and Qwen3.8-Max lost their points [12]. The post's diagnosis is blunt: "That is a training-data artifact with a version number on it." [13] It is a reasonable guess. The harness scores predicted exit codes and last lines [6]. It has no view of what any model was trained on.

Against the floor, the two family scores split. Qwen3.8-Max's 0.40 is 0.046 below the whole-dataset 0.446, and DeepSeek's 0.52 is 0.074 above it [3]. The post describes both as "barely above the 0.446 floor" [23]. It does not say how many of the 19 shared cases are set -u cases, or what a stub scores on that family alone.

The harness bug came from the interpreter's own output format. Bash 3.2 prefixes its error messages with the path the script was invoked with [15]. In case e67 the measured stderr was "scripts/e67_setu_undefined_scalar.sh: line 6: undefined_scalar: unbound variable" [14]. All four models predicted exit 1, and none produced that line [14]. The prompt had said only "Invocation: ... /bin/bash <file>" [15]. The affected cases were grading a guess at the author's directory layout [15]. If no model could make that guess, the best possible score on the full set was 1 - (6 x 0.4) / 57, about 0.958 [2]. The fix put script_path in the prompt and left the measured labels untouched [16].

The method transfers better than the scores. The ground truth is whatever env -i PATH=/usr/bin:/bin /bin/bash printed on macOS 27.0.1 [4]. The author's prompt was wrong about what bash 3.2 would print, and the captured output is what showed it [15]. For the model numbers to carry over to reviewing real scripts, those scripts would need to look like these 57: small cases that each isolate one difference, read closed-book [4] [10]. The claim that the differences "bite production scripts" is the post's own [18]. According to the post, an autonomous agent working for MonkeyRun ran every command [17], and the scripts, measurements and harness are published for anyone to rerun [20].

What to watch

  • Full 57-case runs for Kimi, DeepSeek and Qwen3.8-Max, the only basis on which the author says a ranking would hold.
  • Rescored results after the script_path fix, showing whether the six path-prefixed cases now reach full marks.
  • A stub floor computed for the set -u family alone, to settle whether Qwen3.8-Max's 0.40 is worse than a stub on those cases.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories