Build1 publisherNot yet confirmed elsewhere3 min readPublished
A refactoring benchmark stops the best agent at 41.2%, and the tests are the story
SWE-Bench ProMax puts frontier coding agents on 170 curated refactoring commits. The number that should move procurement is a different one: nearly 60% of unsolved SWE-bench Verified tasks have flawed tests.
The Engineer · Build desk
What happened
- AI coding agents still struggle with large-scale refactoring, with the best model achieving only a 41.2% resolve rate on a new refactoring-focused benchmark developed by researchers at Shanghai Jiao Tong University, Peking University, Douyin Group, and other institutions.
- SWE-Bench ProMax is a multilingual code refactoring benchmark of 170 instances drawn from real commits.
- The benchmark spans seven programming languages: Python, Java, TypeScript, Go, C, C++, and Rust.
- A recent audit cited by the researchers shows nearly 60% of unsolved SWE-bench Verified instances contain flawed tests.
- A 41.2% resolve rate on 170 instances corresponds to roughly 70 instances resolved and roughly 100 unresolved.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Researchers at Shanghai Jiao Tong University, Peking University, Douyin Group and other institutions have published SWE-Bench ProMax, a multilingual refactoring benchmark of 170 instances taken from real commits across seven languages, and the best model managed a 41.2% resolve rate on it [1][5][6]. If you are buying agents on the strength of a SWE-bench score, the more consequential line in the writeup is the audit the benchmark's authors cite: nearly 60% of unsolved SWE-bench Verified instances contain flawed tests [2].
Convert the headline figure into work and it gets concrete. At 41.2% of 170 instances, the best model clears roughly 70 tasks and leaves about 100 on the table [14]. The instances are not scraped and shipped: the creators say every one went through multi-stage curation, with rewritten issue descriptions, manual review of test suites to strip out tests that were too narrow or too broad, and removal of tasks that lacked complexity or cross-file scope [3]. That last filter is the design choice that matters, because it is what makes the set a refactoring benchmark rather than a bug-fix benchmark. The researchers describe the result as a meaningful and unsaturated challenge for current agents [7].
Why refactoring in particular. Shane Warden, principal architect at ActiveState, told The New Stack that strict refactoring "demands zero tolerance for error, zero tolerance for behavior changes, and complete reversibility" [8]. His objection to the prevailing method is structural, not stylistic: engineers who treat models as text-processing engines feed a whole codebase into a large context window, prompt, and wait for multi-file diffs [9], an approach that assumes deep understanding of large systems. "I don't believe that premise," Warden said. "I believe that token proximity does not guarantee structural understanding" [10].
Vojtech Pavlik, senior director of technical strategy for core infrastructure at SUSE, put the ceiling more bluntly, saying there are no LLMs that can read a large codebase and understand it all at once, and that the capability is "still far out of reach" [11]. He added a mechanical reason: very large models need highly optimized attention algorithms such as DeepSeek Sparse Attention to be usable at all, and on a large or convoluted codebase that maxes out the context window, that can mean missing important observations [12]. His second failure mode is the one that does not show up in a resolve rate at all. Where code meets time, he said, you get race conditions, lost idempotency, lost atomicity and incorrect retry handling: code that passes all tests and works if the user follows the spec exactly [13].
The procurement problem sits on top of all this. Evaluation quality for coding agents is considered to be in decline, partly because frontier models can sail through benchmarks whose solutions have leaked into training sets [4].
What to watch: whether SWE-Bench ProMax stays unsaturated once vendors start targeting it, and whether anyone publishes per-language results. The source does not give the language split, and 170 instances over seven languages averages about 24 each, which is thin ground for claiming an agent is good at Rust or C++ [15].