Build1 publisher3 min readPublished
The three possible outcomes were written down before the 380 runs began, which is why a mostly unfavourable result still carries information: aggregation cut one summary by 83 percent while every timed task got slower.
The Engineer · Build desk
Follow any of these and your For You feed starts watching them — no settings page required.
Compiled by The EngineerSomething wrong?How this is made
Ground truth was recomputed before every run of the tuned corpus, and the command that got scored was pulled from the transcript rather than from the agent's own account of what it had done [16]. That second choice is the one worth stealing. An agent that claims a Nushell pipeline and then quietly runs something else is a failure the author already hit: in that environment Claude Code substituted some commands with different implementations [9].
The stated design also reconstructs its own totals. Four tasks, two variants, 25 repetitions per variant gives the 200 microbenchmark runs [11][1]. Five task families, ten repetitions per family per arm, two arms gives 100 [15][2]. Add the 50 held-out runs and the 30 observations from the aggregate case and the headline 380 closes exactly [5][3]. Multiplication is the cheapest audit available on someone else's benchmark, and plenty of tables fail it.
Nushell came out slower on every timed task, and it did not always reduce context [12]. The place it paid was real aggregation, where a container summary fell from 3,353 bytes to 566, a saving of 2,787 bytes or about 83 percent [13][4]. The semantic case is separate and survives the timing loss: comparing sizes and dates as typed values carries fewer implicit assumptions than text columns you manufactured and then decided how to sort [2][14]. For the byte figure to transfer, your agent's task has to be an aggregation whose text form is mostly repeated structure, and the saving has to be worth a slower run. Nothing here replaces zsh; Nushell is a selective route with narrow activation rules [3][8]. If the task is one command that already works, the author's own policy says the structured route is unnecessary [7].
The wrapper is where that policy is enforced rather than described. nu-query returns JSON, imposes a timeout, and reports whether it truncated rows [8]. The truncation flag matters more than the JSON, because a summary that silently drops rows fails in the same direction as the text pipeline it replaced. nu-query is also deliberately absent from the auto-approved permission list, on the grounds that nu -c can modify the system [8].
The most useful material predates the measurement. In the test tree, ls **/* found roughly 46,400 files while ls -a **/* found about 127,900, so the glob without -a was hiding some 81,500 paths, near 64 percent [10][5]. The author presents this as a dated observation about that environment rather than a property of Nushell [10]. Any file comparison run before it surfaced was scoring two different trees.
Which is why the arm that would settle the pre-registration is the one still to read. The three outcomes were defined in terms of accuracy against cost [4], and the accuracy figures for the 100-run corpus are incomplete in the material available, breaking off at "rose from 23/", with no numbers given for the held-out or aggregate arms [18]. The author's own summary is bounded exploratory evidence, with part of the integration tuned during the process [6]. The reusable part is the harness: a named regression case, two of five families where the tool should be refused [15], ground truth recomputed per run, commands read from transcripts [16]. The timings belong to one file tree, in Spanish, on a Sonnet-family model [15][16].
Ranked by verification strength, evidence, and original report placement.
Text shell pipelines depend on text, column positions, and options whose behaviour can differ across implementations, while Nushell preserves tables and typed values such as dates, numbers and file sizes throughout the pipeline.
The author did not want to replace zsh and used Nushell as a selective route instead.
The author states the runs are a large number of repetitions across only a few task families, that part of the integration was tuned during the process, and that the results are bounded exploratory evidence rather than a universal test.
The routing policy uses the least complex tool that can solve the task robustly, and the practical rule is that if Nushell merely runs, inside another shell, a command that already works well, it is unnecessary.
In the tested tree ls **/* found roughly 46,400 files while ls -a **/* found about 127,900, so omitting -a excluded close to 64 percent; the author describes this as a dated warning about a silent failure observed in that environment, not a general property of Nushell.
Nushell was slower on every timed task and did not always reduce context.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One author's numbers, internally consistent
Everything rests on one developer's machine and one post. What lifts it above anecdote is that the counts reconcile: four tasks by two variants by 25 repetitions is the 200 he reports, five families across two arms at ten repetitions is the 100, and the arms sum to 380. He also volunteers the weakness a promoter would bury, telling readers the three positive families were reused during tuning and that his Fisher p of 0.0105 therefore describes the corpus rather than confirming anything. Against that: no replication, no second observer, and 80 of the 380 runs report no result in the text we have.
One workstation, no second user
Use here is one person deep. The skill, the activation rules and the nu-query wrapper run in the author's own setup, and the post names no other developer, team or repository that has picked any of it up. Nushell and Claude Code are both widely used software, but this reporting measures neither's install base; what it measures is a routing policy that exists on one machine and is deliberately fenced off from auto-approved permissions.
Findings undersold
The result works against the author's interest in an exciting one: slower on every timed task and no dependable context saving, with a single 83 percent win where aggregation was genuine. He downgrades his own significant p-value and labels the whole exercise bounded exploratory evidence. If anything is misweighted, it runs the other way — the two findings most likely to save another developer a day, a glob that silently skipped close to 64 percent of files and a nu -c call that returned null with exit code 0 before a 77-second search, sit below a methodology walkthrough instead of leading it.
Personal-brand publishing, no vendor stake
dev.to hosts this, but it is self-published personal work, illustrated from the author's own site and prompted by a podcast episode from Lorenzo Carbonell of atareao.es. No vendor is paying for a favourable Nushell verdict, and none is delivered. The remaining pull is reputational: methodology posts reward a clean positive finding, and this one reports mostly the opposite while flagging that the corpus was tuned by the same hand that measured it.
Solid on method, blind on two arms
Middling, for a specific reason. The design, the harness and the two reported arms are described in enough detail to assess, and their internal arithmetic holds. But our copy of the post ends mid-sentence after the A/B section, so we can say nothing about the held-out tasks or the thesis-reconstruction case — exactly the arms that would have tested whether the tuned-corpus gain survives outside the corpus it was tuned on.
build
255 tools, 71,929 tokens: the standing charge hidden in your MCP config1 publisher
build
A Retention Policy for Agent Memory: Flag Unused Skills at 30 Days, Archive at 901 publisher
build
52 days of zeros: what a cost hook records when the payload never had the numbers1 publisher
leadership
Anthropic and GitHub have moved AI costs from the seat to the meter1 publisher
Publishers with included, body-backed reporting in this cluster.
1 article · September 6, 2026