Skip to content

Build1 publisher3 min readPublished

2,513 tool calls, zero refactorings: what agents actually do when you ask them to refactor

JetBrains traced a model through fifteen C# refactoring tasks and found it simulating structure with sed, git and the compiler. Wiring in Rider's real engine cut median time by 83%.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying 2,513 tool calls, zero refactorings: what agents actually do when you ask them to refactor
Generated illustration

What happened

  • JetBrains traced a frontier model through fifteen C# refactoring tasks; across 2,513 tool calls it performed a structural refactoring operation exactly zero times.
  • The agent performed zero structural refactorings not because it was avoiding them, but because it had none to call.
  • In the baseline trace the model piped text into interactive commands 468 times, called git 422 times and sed 392 times.
  • In the baseline trace the model ran dotnet build 163 times.
  • JetBrains concluded the agent was not compiling to check finished work but compiling to find out what its last edit had done.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

An agent asked to refactor C# does not refactor. JetBrains traced a frontier model through fifteen C# refactoring tasks and counted 2,513 tool calls containing exactly zero structural refactoring operations [1], not because the model avoided them but because it had none to call [2].

What it reached for instead: piping text into interactive commands 468 times, git 422 times, sed 392 times [3], plus 163 invocations of `dotnet build` [4]. Those three text-shuffling habits alone account for roughly 51% of the trace [1].

The build count is the interesting one. JetBrains' reading is that the agent was not compiling to check finished work, it was compiling to find out what its last edit had done [5]. A correct rename has to follow overload resolution, partial classes, explicit interface implementations and documentation references, and tell a type called `Order` from the word "order" in a comment [6]. None of that is recoverable from a regular expression, so the agent guesses in text and lets the build score the guess [7]. Rider's engine, powered by ReSharper, works from a resolved syntax tree that already knows which declaration each identifier binds to, which overload each call resolves to, and where every reference lives across the solution [8].

As of 2026.2.1, Rider bundles a skill called `refactoring-code` that lets an agent invoke those refactorings rather than approximate them; it ships with the IDE, needs nothing switched on, and activates when an agent is asked to refactor C# [9].

The second arm of the evaluation is where the money is. Median task duration fell from 157.9 seconds to 26.6 seconds [10], an 83% reduction [2]. The 95th percentile fell from 346.4 seconds to 56.9 seconds [11], because the slowest runs were the ones stuck in the edit-build-read-error cycle and those runs stop existing [12]. That puts the new p95 at roughly a third of the old median [3]. `dotnet build` dropped from 163 calls to 3 [14], a 98% cut [6], and total tool calls fell from 2,513 to 926 [15], about 63% fewer [4].

Note what did not happen: the agent kept editing text, sed remains its most-used tool, and the eight refactoring operations account for only 167 of those 926 calls [16], under a fifth [5]. The change is division of labour, with ordinary edits staying in the editor's medium and structural changes going to the engine [17].

The caveats are visible in the methodology. Both arms ran gpt-5.5 through the Codex CLI, roughly ten times per task, with skill availability as the only difference, compared using a paired permutation test [18]. The eight operations were picked for clean contracts, meaning a defined target, a defined result, and a refusal when the change is unsafe [19], covering rename, extract method, extract interface, extract base class, change API signature, move type to namespace, reorganize namespaces and safe delete [20]. Fifteen tasks spread across those eight, most in an easy and a harder variant [21]. This is a vendor measuring its own product, and the three published headline figures are time, cost per solved task and tool calls [13]; "per solved task" implies some runs did not solve, and the solve rates are not in the published numbers [13].

Worth watching: the same release bundles a `debugging-code` skill that lets agents set breakpoints and inspect values across C#, F#, C++ and mixed projects [22], extends quality-check hooks to Codex alongside Claude Code [23], and flips ReSharper to out-of-process by default [24]. The pattern to track is whether other toolchain vendors expose their resolved-model engines to agents at all. The number that will matter to operators is not the 83% but the refusal path, because an engine that declines an unsafe delete is the part sed can never imitate [19].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories