Build1 distinct publisher3 min readUpdated
A dev.to writer's Arkanoid bake-off is six agent sessions and one written spec. Any team can copy it this week, as long as nobody trusts the peer rankings.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Six agent sessions and one written spec. That is the entire apparatus: three builds from identical instructions, then three judging passes over the same anonymised trio [1][4][3]. No eval framework, no leaderboard.
The judging half is the part worth copying and the part worth distrusting. Two of the three judges put their own entry first without knowing it was theirs [10], a 2-in-3 self-preference rate across three votes [1], which leaves exactly one cross-vote in the whole run: Gemini's, for Claude [4]. As a ranking, that is unusable. As a control, it earns its keep. An agent asked to grade work it produced is, on this evidence, not a neutral reviewer, and that applies to the review step in your pipeline as much as to a contest.
The build artifacts are where the run actually pays. Codex wrote tooling to check that source files stayed under the maximum line count the specification asked for [8], which means a number in the spec became a mechanical test rather than a hope. Claude spent its session differently: 12 stages, procedural audio, a multi-phase DOH boss, accessibility and responsive work [7], reached through a loop of running tests, launching the game, screenshotting it and inspecting its own screenshots [12], and substantially more tokens than the other two [6]. Gemini delivered a working game with real Arkanoid functionality and much less of the verification and architecture [9][11].
The gap nobody measured is the one that shows up on an invoice. Claude's session ran past 20 minutes against Gemini's rough 5, at least a fourfold spread [2], and the author is straightforward that there was no stopwatch and these are observations rather than benchmark figures [5]. If you rerun this in-house, wall clock and token count per build are two columns that cost nothing to log and are the only ones the original could not report. They are also the numbers that decide whether the thorough agent is worth pointing at a small ticket.
The design choice that makes the whole thing legible is the refusal to help: no fixing mistakes, no "you forgot this feature", no second pass, so whatever the agent called finished was its submission [3]. That converts the exercise from a test of prompting stamina into a measurement of each agent's own threshold for done, which is what lands in a pull request on a normal Tuesday. The task selection follows the same logic, small enough for one session and broad enough to expose physics, architecture, UI, audio, controls and testing decisions [13].
With one task and one attempt each, nothing here settles which model is better. It settles that a comparison you can act on costs a written spec and an afternoon, and that the scoring half needs a judge with no entry in the contest.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The three agents were Claude (Opus 5), Codex (GPT 5.6 Sol) and Gemini (Gemini 3.7 Flash), all set to medium.
All three games were given back to all three agents, anonymised as CL, CO and GE; judges were told they were ranking three contest submissions by creator initials and were not told who built which, nor that the entries were AI-made.
Claude ranked CL first and Codex ranked CO first, each unknowingly voting for its own entry, while Gemini voted Claude's entry into first place.
The dev.to author gave three coding agents the exact same task: build an Arkanoid-style browser game from the same specification.
The author did no fixing of mistakes afterward, gave no feature reminders and no extra passes; whatever each agent decided was finished was its submission.
Rough runtimes were Gemini about 5 minutes, Codex about 10 minutes and Claude more than 20 minutes; the author says he was not timing with a stopwatch and calls these rough observations rather than benchmark numbers.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single self-reported run with inspectable artifacts but no instrumentation
Everything rests on one first-person dev.to post. The strongest evidence is concrete and checkable — three hosted playable builds and specific architectural detail per entry — but the quantitative spine is absent: runtimes are explicitly non-benchmark estimates, token usage is asserted with no figures, the judging rubric and results table are not reproduced, and the model version labels are unverified. No replication, no repositories, no second publisher.
One hobbyist run, artifacts deployed, no external uptake
Adoption evidence is limited to the author's own activity: six agent sessions and three deployed playable builds. Nothing in the supplied material shows another team running this harness, any organisation adopting agent-vs-agent judging, or any vendor or tooling response. That is real but minimal traction.
Headline generalisation outruns three judgments
Positive gap: the framing invites a general conclusion about agents voting for themselves and about relative agent speed, while the underlying data is three judgments in one unreplicated run with uninstrumented timings and unverified model labels. The gap is moderate rather than severe because the author hedges heavily himself — disclaiming benchmark status, rejecting the 'narcissist' reading, and volunteering the screenshot round in which Claude ranked Codex's visuals first.
Personal-brand and traffic incentives, no disclosed vendor stake
The author is an independent practitioner publishing on a developer-engagement platform, and the three demo builds are hosted on subdomains of his own site, giving a clear attention and personal-brand incentive to produce a surprising result. Nothing in the supplied source indicates payment, sponsorship or a stake in any of the three model vendors, and the self-limiting caveats cut against maximal-engagement framing, so the incentive load is moderate rather than high.
Confident about what was done, not about what it means
Confidence is moderate: the procedure and the artifacts are described clearly enough and the builds are inspectable, so the account of what happened is credible. Confidence in any generalisation is low — one publisher, one run, three judgments, no instrumentation, no third-party verification of the model versions or codebases.
build
Claude's system prompt grew ninefold in two years. Version yours like code.1 distinct publisher
build
Claude's prompt cache dies quietly in agent loops: the 20-block lookback nobody configures1 distinct publisher
build
Four frontier models in four days, and the cheapest number in your agent plan has an expiry date1 distinct publisher
build
Safety fixes ship in new model versions. The regression stays with whoever pinned the old one.1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 22, 2026