Build1 distinct publisher3 min readUpdated
One good demo is not evidence that a new prompt, skill or rules file helped. A control arm, thirty tasks and a held-constant model will tell you, within limits worth knowing before you build it.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Run the arithmetic on the recipe before building it. Two arms of ten to thirty tasks is twenty to sixty agent runs per rule change [4], which is cheap enough to do between other work. It is also small. Thirty tasks scored pass or fail resolve nothing finer than one task, about three and a third percentage points [1]. The guide's own cautionary example, a standard that lifts quality by 3% while doubling token use [10], sits under that floor: 3% of thirty tasks is nine tenths of one task [2]. The doubled token count you will see, because it is counted rather than judged. The quality gain you will not, and if you adjust the rubric until you do, you have rebuilt the evidence-free rule change the harness was meant to replace [2].
So point it at delivery instead. The failure the guide calls unreliable instruction delivery is the one that shows up at this sample size: the wrong instruction file loads, the rule sits too deep in context to matter, a stale example gets over-followed, or no skill is selected at all [4]. A migration safety skill that never loads because the ticket said "fix signup bug" is not a 3% effect [6]. It is a count, and counts survive twenty runs in a way rubric averages do not, which is why the guide tells you to measure selection reliability and not only output quality [13].
Conflicts are the second thing worth a fixture. When one file says prefer fast minimal changes, another says always add complete tests, and a third says avoid touching test snapshots, precedence is being decided at runtime, per task, by the model, and nobody on the team has written down which rule wins [8]. A conflict case in the task set turns that into a reviewable artefact, which is the actual argument for treating standards like code: versioned, reviewed, tested, rolled out [14].
The cost side needs a verdict per workflow, not one global answer. Take the guide's own tradeoff at face value: 3% more quality for twice the tokens is roughly 94% more spend per acceptable output [3]. The guide's position is that this can be fine for high-risk work and probably is not fine for every background automation [11]. That means the same variant can pass and fail in the same organisation, and the harness has to record which population it was scored against.
What makes any of this affordable is the control on the runner: same agent, same model family, same tool access [12]. Nothing about model capability is under test, so no ground truth about the model is required. Your prose is under test, which is the question teams actually have when they add a rule [3]. Worth noting that the source stops mid-sentence on the scoring rubric, so the cheapest part of this design is specified and the expensive part is not.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
AI agents can look reliable after one impressive demo and still fail once real users, messy repositories and conflicting instructions are involved.
The guide argues the dangerous part is not that agents make mistakes but that teams change agent rules based on vibes rather than evidence, and that standards need tests beyond model evals and unit tests.
The practical question the guide says a test system must answer is whether a new rule, skill, prompt or tool instruction actually made the agent better.
The guide states the main risk is not only bad instructions but unreliable instruction delivery: an agent may load the wrong instruction file, ignore a rule buried deep in context, over-follow a stale example, or select no skill at all.
An AI agent standard is defined as any reusable instruction that changes how an agent works, including repository rules such as AGENTS.md, CLAUDE.md or Cursor rules, coding guidelines, skill descriptions, tool usage policies, review requirements, support response rules, RAG grounding rules, approval policies and prompt templates used across tenants.
Example of skill non-selection: a team creates a database migration safety skill, the agent edits a migration file, but never loads the skill because the task was worded as "fix signup bug".
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single self-published guide, no results reported
All claims trace to one dev.to practitioner post. The methodology is internally coherent and specific, but nothing in it is empirically demonstrated: no run of the harness, no task-set artefacts, no pass rates, no worked control/variant comparison, and the '3% quality, double tokens' figure is an illustration rather than a measurement. The only quantitative work available is arithmetic derived from the guide's own stated numbers, which cuts against the design rather than validating it.
No adoption signal in supplied sources
The cluster contains no release, deployment, benchmark, usage disclosure, pricing or licensing event. The source names artefact types (AGENTS.md, CLAUDE.md, Cursor rules) but reports no team, product or repository that has adopted the described experiment system, so adoption cannot be scored without inventing facts.
Modest tone, but implied precision exceeds the design
The article is deliberately unhyped — it disclaims any vendor pitch or magic framework and says the goal is not perfect science. The overstatement is narrower and quantitative: it presents a 10-to-30-task A/B as sufficient to answer whether a rule change helped, while its own 3% example effect is smaller than the one-task (about 3.3 point) resolution of a 30-task pass/fail arm, and it omits repeat runs, grader variance and the 20-to-60-run cost per change. Positive but small, because the prescription is directionally sound and the gap is calibration rather than substance.
Low commercial pull, developer-platform authorship
The post explicitly disclaims a vendor pitch and promotes no product, framework, pricing tier or license; the artefacts it names are generic and vendor-plural. Residual incentive is the ordinary one for self-published developer-platform writing — audience and credibility building through prescriptive methodology — and the source discloses no employer or affiliation, so it cannot be fully cleared.
Clear on what was said, thin on whether it works
What the source asserts is unambiguous and directly quotable, and the derived arithmetic follows deterministically from its own figures, which supports moderate confidence in the read. Confidence is capped by there being one publisher, zero adoption or outcome data, a truncated article body, and no external corroboration of either the failure taxonomy or the harness's effectiveness.
build
A "Done." is a claim about the world, not a sentence you can grade1 distinct publisher
build
The linter that passed everyone who ignored it and warned everyone who complied1 distinct publisher
build
The Context Tax: Your Developers Are Doing Unpaid Platform Work Every Session1 distinct publisher
build
Deleting a 1,350-line CLAUDE.md broke two rules out of twenty1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 23, 2026