Skip to content

Build2 publishers3 min readPublished

Frame-hash replays let CodeScene's agents refactor 300,000 lines of Street Fighter III for $4,000

CodeScene's agents refactored 300,000 lines of Street Fighter III in three weeks for about $4,000 in tokens, taking its Code Health score to 10.0. The run relied on a frame-by-frame replay check and on the score the agents were told to optimize, so the promised savings on later feature work still need their own measurement.

The Engineer · Build desk

Photograph accompanying Frame-hash replays let CodeScene's agents refactor 300,000 lines of Street Fighter III for $4,000
Photo: codescene.com

What happened

  • CodeScene ran the project with researcher Marcus Borg on a decompiled version of the 25-year-old game, written in C.
  • Along the way the agents wrote their own playbook of 22 recipes and 82 notes, including codebase-specific ones such as Shared Index Range and Uniform Step Table.
  • The team settled on Claude Code with Opus after finding it significantly better than Codex with Sol at documenting patterns, while smaller models stalled on files.
  • Daniel Webb, one of the two engineers on the project, said the work was merged to main on a fork through 54 pull requests.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint A team whose code has no per-change behavioral check as strict as frame-hash replay would be running the same agents without the gate that made these changes safe to merge.
  • cost Reviewers pay the part of the bill that the token figure leaves out: 54 pull requests averaging about 54 commits and 4,700 changed lines each.
  • contradiction The podcast notes put the starting score at 5.4 and InfoQ at 5.6, so even the baseline of the project's main evidence of improvement differs between the two accounts.

Street Fighter III came with its own regression oracle. The game already hashed each rendered frame and compared the hashes [6]. CodeScene's replay-trace harness compared the rollback state hash frame by frame, so behavior could be checked after every change [12]. Every agent edit got a pass or a fail inside the loop. Paolo Perrone wrote: "most refactor claims i've read rest on a green test suite, which only tells you the tests survived. replaying traces against a fighting game sets a much higher bar." [16]

The engineers who built the oracle acknowledge a gap in it. Webb, who is CTO at NeoSee, said the harness ran as a pre-commit hook and that some failures were fixed without being observed [23]. Asked whether it caught subtle frame-timing regressions, he said there might be no definitive answer [23]. Performance is also an open question. Marc Bouvier asked about framerate, memory use and input latency, and Webb said a performance specialist was being brought in [22].

The quality signal ran through the same loop. A CodeHealth MCP server gave the agents a deterministic score to optimize, and they used it to judge whether a transformation had helped [11]. The finishing 10.0 is that same score [10]. Tornhill, CodeScene's founder, calls 9.5 or higher "AI-optimal code" [1][4]. Denis Baltor argued that DRY concerns duplicated knowledge and intent, and that recipes built on loops differing only in their ranges may be treating lines that look alike as shared knowledge [19]. Asko Nõmm said architecture goes unmeasured, so code can look healthy while fundamental problems surface later [21]. One choice deserves credit. The playbook logged failed attempts, including transformations that made Code Health worse [14].

About $4,000 in tokens bought 2,903 commits and 252,055 modified lines [8][9]. That works out to roughly $1.38 a commit and 1.6 cents a modified line [2][3]. At those rates the invoice is the least interesting thing the project produced. The agents touched about 84% of the 300,000 lines [1]. Tornhill said: "I've never ever seen anything close to this before. A rewrite at that scale would have been a 12-to-18-month project. And now that can be automated for a fraction of the cost at a fraction of the time." [5]

The "few days" in the podcast notes covers the uplift alone. Preparation took about two weeks as a side project, while the team tried different models and patterns [7]. Two weeks plus a few days comes to roughly the three weeks InfoQ reports [6].

I think the refactor result holds within its harness. The payoff is a separate claim. The podcast notes say the team measured how the cleanup led to better and cheaper AI development of new features [2], but the public show notes and InfoQ's report do not include those figures. Nõmm raised a second problem. Claude Code and Codex are tuned to their own models, he argued, so it is unclear how much of a result measures the model and how much measures the harness, and cost would vary the same way [20]. Konrad Otrębski, a tech lead and consultant, wrote: "I think the real true test of AI capability here would be to offer this courtesy of refactoring to some famous open source project, say Grafana. The definition of done is ofc merging it to master." [18]

What to watch

  • Release of the feature-development measurements the podcast says CodeScene took, with cost per feature before and after the cleanup.
  • The performance specialist's results on framerate, memory use and input latency for the refactored game.
  • An agent-driven refactor merged to master of a maintained open-source project such as Grafana, the test Otrębski proposed.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories