Build1 publisher3 min readPublished Updated
A 755-line AGENTS.md moved one of 26 assertions in a controlled agent test
Eval harness agents-md-evals found 25 of 26 assertions passed identically with or without a 755-line AGENTS.md file. The finding rests on that one file, and it assumes the agent loaded the file in the first place.
The Engineer · Build desk

What happened
- The harness credits a codebase-teaches-patterns effect, where package.json, existing imports and the directory tree show the model a project's stack before it reads any rules.
- For well-structured projects, the harness puts the realistic improvement an instruction file delivers at 3 to 10 percent.
- Control runs follow a seven-phase isolation protocol so the instruction file is genuinely absent from the agent's context.
- In Codex, according to a writeup the post cites, only one instruction file per directory is read, and an AGENTS.override.md replaces the AGENTS.md beside it.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost Rules that fail the A/B are still sent as context on every request, so a mostly redundant file costs tokens on every call for guidance the model would have followed anyway.
- decision Rule-writing effort moves to what code cannot show: why two services must ship together, workflow steps no file records, tie-breaks between two sound choices, and domain knowledge held by people.
- constraint A fair control needs the rules file physically absent, because Claude Code and similar agents auto-load it into every conversation and subagent, even when told to ignore it.
- exposure Teams can mistake a silent Codex loading failure for weak instruction-following and end up rewriting rules the model never saw.
The 25-of-26 figure comes from testing a single 755-line file against 26 assertions [1]. The harness is small and built on Anthropic's skill-creator framework [5]. The dev.to post that reports it does not name the repository, the model, or how many runs produced the result.
For the number to hold on another team's project, their file has to look like the tested one. According to the post, lines such as "use TypeScript strict mode" or "follow functional patterns where possible" teach a model nothing it could not infer from opening two files [20]. The result covers a file made mostly of lines like those, in a repository that shows its conventions just as plainly. The assertions matter as well. The harness writes its eval prompts from the project's own git history [9], so the test measures the work that repository actually does.
Passing identically on 25 of 26 is about 96 percent [1]. The README's general claim is lower: most instruction files, it says, are 80 to 95 percent redundant [2].
I think the isolation work is the most careful part of the project. The harness moves the files to /tmp/ because, according to the post, a rename on the same path sometimes still resolves [8]. So a control run that only renamed the file could still load it. The keep test is simpler. A rule stays if it meets at least one of six criteria: specificity, behavior change, surprise, pain, frequency, or no feedback loop [10]. If it fails all six and fails the A/B, it is removed [10]. The post's author wrote: "Everything else is usually the model guessing correctly and your file taking credit for it." [19]
All of this assumes the file reached context. For Codex, a writeup the post cites lists several ways it does not, and none of them raises an error [18]. The instruction search stops at the working directory. A rules file deeper in the tree is invisible when the agent starts at the repo root [14]. Codex builds the instruction chain once per run, or once per launched session in the TUI, and a running session does not see edits [17]. The docs annotate a file in their own sample tree as "Ignored because an override exists." [13]
The size cap is the one I would check first. Per the same writeup, the `project_doc_max_bytes` setting defaults to 32 KiB, and two doc pages disagree on whether it applies per file or to the whole chain [15]. Spread over 755 lines, 32 KiB is about 43 bytes a line [2]. A file as long as the tested one, with a longer average line, would go over the default under either reading [2]. The post's author wrote: "The reliable move is to raise it and verify." [16]
What to watch
- Whether agents-md-evals publishes the repository, model and run counts behind the 25-of-26 result, or results from repositories whose conventions are less visible in code.
- Whether the Codex docs settle if the 32 KiB project_doc_max_bytes default counts one file or the whole instruction chain.
- Whether Codex starts warning when an override, the working-directory boundary or the size cap keeps a rules file out of context.