Build1 distinct publisher3 min readPublished
A preregistered pilot burned 30 agent attempts and both arms passed everything, a result that describes the three fixtures better than it describes the skill. The author reports the ceiling as his finding.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Start with the interval, since it does the real work in this pilot. Newcombe put the treatment-minus-baseline difference at 0.000000, with a 95% interval of [-0.203883, 0.203883] [3]. That window is about 41 percentage points wide [1]. Scale the half-width by the 15 attempts in an arm and you get 3.06 attempts [2]. So the design cannot exclude an effect worth three passes out of fifteen, in either direction. Two identical perfect scores are easy to mistake for a finding when they are not one.
Part of that comes from the scoring rule. An attempt passes only as the logical AND of eight checks: artifact contract, source integrity, visible and hidden sensitive-value removal, secret and path removal, preservation of safe content, byte-exact Markdown, and byte-exact findings JSON [6]. As a release gate, that is the right shape. As a measuring instrument, it discards everything except the final bit. Once the baseline clears all eight on every attempt [1], there is no margin left to record, and the finest difference an arm can express is one attempt, or 6.7 points [3].
Two conditions would have to hold before that 15/15 said anything about your agents. Your task family has to fail at the baseline often enough for 15 attempts to see it. And your agents have to be cued: both arms got the same generic instruction to identify and read applicable skills, so what was tested was encouraged availability rather than spontaneous discovery [8]. Two further limits are the author's own. The inputs were reserved synthetic examples, and he says the result should not be read as evidence of real-world PII safety [9]. The intervals describe repeated runs on three fixed fixtures and do not estimate generalization across the broader population of public-release tasks [16].
The cheap fix is ordering:
1. Run the baseline arm alone against candidate fixtures. 2. If it saturates, harden or replace fixtures until it fails at a rate 15 attempts can resolve. 3. Then freeze the protocol and spend the treatment arm.
The harness deserved a better benchmark than the one it got. Protocol, runner, skill, fixtures, order and analysis rule were all frozen before the scored batch, and a context-isolated adversarial review had to return GO before the one-shot runner could start [7]. All 30 trials came back healthy, with no diagnostic failures, infrastructure errors, retries or replacement attempts [10]. Most usefully, the run splits one question into three: whether the procedure was available, whether it was used in the intended order, and whether using it actually improved the outcome [17]. The first two were answered.
What stays unmeasured is the part the reviewers flagged: creating, testing, loading, selecting, updating and retiring the skill [13]. The pilot prices none of it, and the author declines any break-even claim because human design and maintenance time was never measured [15]. My reading, which I would revise given a discriminating fixture set: an internal agent eval needs a published baseline failure rate before it can claim to measure anything. This one at least says so, in the author's terms, that the overhead of the procedure is observable while its outcome value remained unmeasured [19].
Ranked by verification strength, evidence, and original report placement.
In the pilot, the baseline arm passed 15 out of 15 attempts.
The treatment arm, which had the frozen sanitizer skill available, also passed 15 out of 15 attempts.
The treatment-minus-baseline difference was 0.000000, with a Newcombe 95% interval of [-0.203883, 0.203883].
The preregistered decision was no clear difference; the author notes that with only 15 attempts per arm the data remain compatible with practically meaningful benefit or harm, and that a 15/15 tie does not establish equivalence.
The design used three unseen synthetic fixtures, baseline and treatment arms, five attempts per fixture and arm for 30 attempts total, a fresh container and fresh agent session for every attempt, and no retries, replacement attempts or early stopping.
One exact binary pass was defined as the logical AND of eight checks covering the artifact contract, source integrity, visible and hidden sensitive-value removal, secret and path removal, preservation of safe content, byte-exact Markdown, and byte-exact findings JSON.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · September 2, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
The third answer: a dead-code tool allowed to say "not traced yet"1 distinct publisher
build
A SKILL.md layer quietly rerouted an agent off the MCP tools it was given1 distinct publisher
build
Three manual interventions in a month, and every guard was working as designed1 distinct publisher
build
Six MariaDB versions, one real difference: the only reason to leave 10.6 is the July 2026 clock1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Disciplined method, single witness
The procedural hygiene is better than most blog-scale experiments get: pre-commitment before scoring, a gating adversarial review, a fresh container and session per attempt, no retries, and a pass criterion tight enough to include byte-exactness. What is absent is anyone but the author — no fixtures, runner or per-attempt logs travel with the numbers, and no second party reran the batch. Thirty attempts across three fixtures then caps how much any of that rigor can conclude: the interval is wide enough to swallow the effect it was built to detect.
One bench, no deployments
The only usage anywhere in this reporting is the author's own: a skill built for the experiment, exercised 30 times against synthetic fixtures inside disposable containers. Nobody else's team, product or pipeline appears. The uptake numbers are real telemetry, but they measure whether one agent read one file in the intended order — not that agent skills are being run anywhere for consequence.
The author argues against himself
A perfect 15/15 in both arms is the easiest number in tech writing to oversell, and the sales pitch never arrives. The tie is called a ceiling effect within four sentences; equivalence is refused; the cost advantage is downgraded to harness estimates with the human labour explicitly missing; the safety reading is fenced off before a reader can reach for it. Our own framing leads with the 41-point window rather than a verdict, which is the honest place to lead. If anything the piece undersells its most durable contribution — the measurement design — by filing it under a failed study.
Credibility, not revenue
Nothing is for sale in this post: no vendor, no tool, no funding, no product whose adoption a null result would help. The pressure that does exist is reputational — a practitioner recovering from a thesis his own reviewers dismantled has reason to look maximally rigorous, and rigor is exactly what the write-up performs. Worth noting where that self-supervision loops: the adversarial GO gate and the three reviewers who broke the original argument were all agents the author himself convened.
Sure what was counted, unsure what it means
We can be fairly confident about the mechanics as reported — the arm sizes, the interval, the read-rate, the exclusions the author names himself are all stated plainly and internally consistent. Confidence drops on transfer: no model or harness is identified, three fixtures cannot speak for public-release tasks generally, and the only verification available is the author's word. So: high trust in the account, low trust in extrapolating from it.