Skip to content

Build1 publisher2 min readPublished

AgSpec nearly doubles speculative draft length by indexing code the way coding agents write it

KAIST and Seoul National University's AgSpec nearly doubles accepted draft length for coding agents by indexing files in the diff and JSON forms agents emit. It lives entirely in the retrieval index and leaves model weights untouched.

The Engineer · Build desk

Illustration accompanying AgSpec nearly doubles speculative draft length by indexing code the way coding agents write it

What happened

  • On SWE-bench Verified, the authors found that standard retrieval indexing misses most of the reusable code already sitting in the agent's workspace.
  • Agents write edits as unified diffs that prefix copied lines with a space or a minus sign, or as JSON tool calls that escape newlines and quotes.
  • The suffix index matches token IDs, so the prefixed or escaped text tokenizes differently and the match breaks at the first character of every line.
  • AgSpec drafts from three corpora in priority order: the live session first, files opened in the workspace second, and shared global references third.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Index memory grows with every file the agent opens: three copies of its text for a diff-based coder, two for a JSON tool-call harness.
  • decision A suffix-index speedup measured on isolated code snippets says little about the speedup inside an agent harness, so teams have to measure it again there before counting on it.
  • constraint Longer drafts pay off most at batch size 1; per the post, at batch size 16 or in multi-agent serving, where verification saturates GPU cores, extra drafted tokens cut serving capacity.

The post's author wrote, "I spent time assuming the problem was cache hit rate or vocabulary drift between sessions" [15]. The paper, arXiv:2610.01108, "points to a much dumber reason," the author wrote [c3, c16]. For a diff coder, the defect is one byte per line [5]. A continuation copied from the raw file can match up to a newline, then fails on the prefix that opens the next line [6]. So a draft copied from disk ends on the line where it started [6]. In escaped JSON the mismatch can also land mid-line, wherever a quote picks up a backslash [5].

AgSpec's fix is a text transform applied when the agent opens a file [8]. The raw file stays indexed for general context. A diff coder also gets one variant with a leading space on every line and one with a leading minus, and a JSON harness gets an escaped copy [8]. The two diff variants match the two kinds of line a diff copies verbatim from the existing file: context lines and deletions [5]. The variants live in the workspace corpus. It starts empty, grows only as the agent opens files, and disappears when the container terminates [10].

I think this is the right first fix for a team already running suffix-index drafting behind a diff-based harness. Retrieval drafting already avoids training a separate draft model [1]. It is small, correct engineering: the matcher compares token IDs, and AgSpec changes the indexed text so those IDs line up with what the model emits [c6, c8].

The corpus split targets a second failure. Earlier systems such as FastCoder leaned on static project datastores [11]. The author wrote that an agent "spending turn four analyzing a traceback needs to draft tokens from the compiler error printed on turn three" [17]. AgSpec's session corpus holds conversation history, tool executions, terminal logs and earlier error traces, and it is discarded when the session ends [10]. The global corpus of standard libraries, framework documentation and common dependencies is compiled once ahead of time and shared across sessions [10].

The gain is counted in tokens: the post gives it as accepted draft length on repository-level edits [9]. It does not report wall-clock latency or throughput. An accepted token saves an entire autoregressive decoding pass, and a rejected one wastes verification compute [13]. For the gain to carry over to another deployment, the agent has to copy long blocks out of files it has opened, emit them as diffs or escaped JSON, and be served at a batch size where rejected drafts stay cheap [c8, c13, c14].

What to watch

  • Wall-clock latency and throughput figures for AgSpec at batch size 16 and in multi-agent serving, set against the draft-length gain.
  • How AgSpec caps or adapts draft length once verification saturates the GPU, the regime the post flags as costly.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories