Skip to content

Build1 publisher3 min readPublished

CodeScene's AI agents refactored Street Fighter III's codebase for $4,000, checked by a frame-by-frame replay harness

CodeScene's coding agents refactored a 300,000-line C codebase in three weeks for about $4,000 in tokens. The run also depended on a deterministic quality score, and the team's own model comparison cannot say how much of the result belongs to Claude Opus.

The Engineer · Build desk

Photograph accompanying CodeScene's AI agents refactored Street Fighter III's codebase for $4,000, checked by a frame-by-frame replay harness
Photo: codescene.com

What happened

  • The run produced 2,903 commits across 726 files and 252,055 modified lines, moving the codebase's Code Health score from 5.6 to 10.0.
  • Along the way the agents built their own playbook of 22 refactoring recipes and 82 notes, including failed transformations that made Code Health worse.
  • Daniel Webb, one of the two engineers on the project, said the changes reached main on a fork through 54 pull requests.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • contradiction The team credits Opus, yet each model ran inside its own vendor's tool, so the case study cannot rank model against scaffolding or pin cost to a model.
  • exposure Teams copying the method inherit a safety gate with an unknown catch rate, because failures fixed inside the hook were never counted.
  • cost The token bill, about $1.38 per commit, is the cheap part; a team without a replay oracle has to build one before the first agent run.

Each change had to pass two checks. The CodeHealth MCP Server returned a deterministic score, so an agent could tell whether a transformation had helped or hurt [5]. A replay-trace harness then compared the game's rollback state hash frame by frame [6]. Daniel Webb, CTO at NeoSee and one of the two engineers on the project, said the harness ran as a pre-commit hook [14]. A failing check blocked the commit until the agent fixed the change.

The codebase is an open-source decompilation of Street Fighter III: 3rd Strike [3]. That choice gave the project an unusually strict behaviour check. "most refactor claims i've read rest on a green test suite, which only tells you the tests survived. replaying traces against a fighting game sets a much higher bar," Paolo Perrone wrote [9]. For the $4,000 figure [1] to transfer, a target codebase needs the same pair of inputs: recorded behaviour it can replay, and a state digest it can compare after every change. Few payroll systems ship with a per-frame state hash.

The unit costs are small. Spread over 2,903 commits [2], $4,000 is about $1.38 a commit [1]. The 252,055 modified lines equal 84% of the stated codebase size [2], though a diff count is not the same as distinct lines touched. The work reached main on a fork through 54 pull requests [10]. That is about 54 commits per request [3].

Code Health rose from 5.6 to 10.0 [2]. It is also the score the agents were optimising [5]. So the final number shows the agents hit the target they were handed. Denis Baltor argued that DRY concerns duplicated knowledge and intent, and that recipes built on loops differing only in their ranges may be collapsing intent into identical lines of code [13]. Asko Nõmm noted that architecture is not measured, so code can look healthy while fundamental problems surface later [12].

The playbook is good engineering. The agents ended with 22 recipes and 82 supporting notes. Some are specific to this codebase, such as Uniform Step Table, which turns heterogeneous calls into table-driven dispatch [7]. They also logged failed attempts, among them transformations that lowered Code Health [7].

The evidence does not show the model mattering less than the scaffolding. The team settled on Claude Opus for most of the work [8]. It reported that Claude Code with Opus was significantly better than Codex with Sol at capturing and documenting the emerging patterns. Files plateaued under smaller models [8]. Nõmm's objection is that each tool is tuned to its own vendor's model. The comparison therefore scores a model-and-harness pair, and cost would vary the same way [12]. What the case study supports is narrower: the team credits the score and the replay check with carrying the work, and inside them Opus went further than the alternatives it tried. Adam Tornhill, CodeScene's founder, wrote that after three decades on large systems this was the first time he had seen what he called "superhuman AI performance at scale" [4].

Asked whether the harness caught subtle frame timing regressions, Webb said there might be no definitive answer, because some failures were fixed without being observed [14]. Marc Bouvier asked whether framerate, memory use and input latency had improved. Webb said a performance specialist was being brought in [15]. Konrad Otrębski, a tech lead and consultant, set a harder bar: "I think the real true test of AI capability here would be to offer this courtesy of refactoring to some famous open source project, say Grafana. The definition of done is ofc merging it to master." [11]

What to watch

  • Results from the performance specialist on framerate, memory use and input latency in the refactored fork.
  • A run that puts Opus and Sol inside the same harness, which would separate the model's contribution from the tool's.
  • Any attempt at Otrębski's test: an agent refactor merged upstream into a large production open-source project.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories