Build1 distinct publisher2 min readPublished
A new Codex benchmark tries to settle multi-agent orchestration by measurement. Its first run qualified the harness and could not report the combined token bill.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The load-bearing idea in this thing is the gap between a job title and a write surface. "Backend agent", "test agent" and "review agent" read like separate concerns, and according to the project's maintainer they will still collide on the same files or sit waiting on each other's stages [9]. The boundary that survives contact is a path: one writer owns `incident/`, another owns `web/`, and neither edits the shared contract [10]. You can check that against a repo before you spawn anything.
The gate itself is six preconditions, and all of them have to hold [1]. Three are facts about the task: two genuine ownership surfaces, exclusive paths per writer, an interface frozen before implementation [7]. The other three are facts about your own discipline: the controller keeps integration, system checks and final review; workers return concise evidence instead of transcripts; a single external acceptance bar judges every execution method [8]. Small change, unsettled interface, or several steps editing the same central files, and you are back to one agent [6].
Now the measurement. The case against extra agents starts with total tokens, duplicated context, handoff delay and integration risk [1]. In the first smoke run, which qualified the harness [21], the environment did not expose aggregate controller-plus-worker tokens [18]. So the harness validated itself on everything except the quantity that decides the question it was built to answer [2]. The maintainer says so plainly: feasibility, not superiority, one candidate per method, overlapping runs, and no valid claim that orchestration was faster, cheaper or more reliable [17]. That is two runs in total [3].
What the run did show is worth separating. The orchestrated candidate came back with live browser evidence across creation, filtering, state transitions, validation feedback, console errors and a mobile viewport [15]. The single agent found and fixed a conflict-message behaviour during final review [16]. Note where that catch came from: the review step that the orchestration recipe insists stays with the controller anyway [8]. Coverage evidence is the honest argument for a second worker [2], and it is an argument about what gets observed, not about speed.
The transferable artifact is the delegation budget, written into the prompt: at most two implementation workers, backend owns `incident/` only, frontend owns `web/` only, and no worker touches TASK.md, the tests, the evaluator files, or the other's paths [19]. That quota costs nothing and it forecloses the failure mode the piece names, where a large goal becomes licence to create an agent per noun in the prompt [20]. The fixture was built with two surfaces on purpose [12]; most tickets have one.
Ranked by verification strength, evidence, and original report placement.
The orchestrated candidate returned live browser evidence covering creation, filtering, state transitions, validation feedback, console errors and a mobile viewport.
The single agent found and corrected a conflict-message behaviour in final review, and an external controller later exercised its live HTTP behaviour.
The result establishes feasibility, not superiority: the runs overlapped and there was only one candidate per method, so it would be invalid to claim orchestration was faster, cheaper or generally more reliable.
The environment did not expose aggregate controller-plus-worker tokens.
The first smoke run qualified the benchmark harness.
Codex How To now includes a dependency-free benchmark for testing that question instead of answering it from intuition.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One self-reported smoke pair
The harness, fixture, evaluator checks and delegation contract are described in verifiable detail and the author publishes a receipt, which lifts evidence above pure assertion. But everything is single-source and maintainer-reported, the comparison is one candidate per method with overlapping runs, the article's own protocol asks for at least three alternating pairs, and the general cost/benefit framing about multi-agent work is argued rather than measured.
Maintainer-side artifact only
Supplied evidence shows the benchmark exists inside the author's own open-source project and has been exercised once. No third-party run, dependent project, deployment or usage disclosure appears anywhere in the cluster, so observed adoption is confined to the maintainer's own repository.
Self-limited relative to framing
The cluster framing suggests measurement settling the multi-agent question, which would be an overstatement, but the source repeatedly pulls its own claims below what a promotional post could assert: feasibility not superiority, invalid to claim faster or cheaper, tokens written as unavailable, and a protocol demanding three alternating pairs before conclusions. The net position is marginally understated rather than inflated, with the residual risk sitting in the headline rather than the body.
Disclosed maintainer stake
The author both maintains and benchmarks Codex How To, supplying the benchmark, the evaluator and the measurements, and publishes on a self-service developer platform with no editorial review. That is a material promotional incentive, partly offset by an explicit disclosure and by self-imposed limits on what the result may be used to claim.
Design credible, results thin
Confidence is moderate-low: the descriptive facts about the fixture, evaluator, gate and delegation budget are internally consistent and directly checkable in the published project, but every empirical statement traces to one publisher, one maintainer and one overlapping pair of runs, with the key cost metric absent and no corroborating source in the cluster.
build
Block's Berd makes a duller argument than its mascots: show the agent's context as product state1 distinct publisher
build
Your reviewing model is reading the diff when it should be reading the session1 distinct publisher
build
Developer habit, priced at $965B: what Anthropic's run actually proves1 distinct publisher
build
Coding agents fail before they compile, and the fix is a sign-off rather than a better model1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 25, 2026