Build1 publisher3 min readPublished
Dropping Claude's parallel tool calls trades a 400 error for wrong answers
A developer's Claude agent crashed on 388 of 1,200 runs because its loop answered only the first of several parallel tool calls. Deleting the extra calls stopped the errors but cut its eval score from 46 to 38 out of 50.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The loop took only the first tool_use block from each response, while 14.2% of tool turns in the logs carried two or more calls, up to six.
- The Messages API returns a 400 when any tool_use id in an assistant turn lacks a matching tool_result in the next user message.
- With the extra calls deleted from history, the model re-requested files it believed it had never asked for, and average turns per run rose to 4.9.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- exposure A first-block loop passes manual tests built on simple questions and fails only when users ask multi-file questions, so the bug ships to production.
- cost Silencing the 400 by editing history turns a logged error into confident wrong answers, and here only a hand-labeled eval set exposed them.
- decision Builders choose between handling several results per turn or disabling parallel calls and paying about 0.8 extra turns per run, which is worth it only when tool order matters.
The per-run crash rate was more than twice the per-turn rate. In the author's dev.to write-up, 388 of 1,200 runs failed, or 32.3% [1], while 14.2% of tool turns carried more than one call [3]. Each run makes several tool turns, and any one of them can be the parallel one. If parallel turns landed independently at 14.2%, two tool turns per run would give a 26.4% chance of hitting one, and three would give 36.8% [3]. A good run takes 3 or 4 round trips [15]. The author wrote that the failure ratio was "too low to look like a broken deploy, too high to ignore" [16].
The day-one loop passed every manual test the author ran [14]. Every one of those tests was a simple question, and simple questions get one tool call at a time [14]. Comparison questions behave differently. Parallel tool use is on by default in the Messages API, according to the author [12]. Asked how two webhook senders handle timeouts, the model requested both files in one turn. The loop ran the first, appended both blocks to history and sent back one result [18]. The next request failed with `tool_use ids were found without tool_result blocks immediately after: toolu_01B...` [17]. The message names the turn index and the orphaned id [17].
The first patch deleted the unanswered tool_use blocks before appending the assistant turn [5]. That satisfies the pairing rule by editing the record. The model's history then said it had asked for one file. Either it asked again, pushing average turns to 4.9, or it answered without the second file [7][8]. Some of those answers described the legacy sender from code the model had never seen [8]. "Editing the assistant's own past turns is gaslighting your agent. It plans based on what it believes it already did," the author wrote [9].
The eval gap is eight questions, or 16 points on a 50-question set [2]. It measures one agent: three tools (read_file, grep, list_dir) over a mid-sized Python monorepo, a Sonnet model, and questions hand-labeled with the correct file and line range [13][6]. I'd expect the direction to hold for any agent where the deleted calls were ones the model needed. The size depends on the workload. It moves with how often a question mix produces multi-call turns, and 14.2% is this mix's rate [3].
The working loop runs every call concurrently with asyncio.gather. It returns all results in one user message and sends failures with is_error: true [10]. The flag matters because a failed call still has an id. Omit its result and the pairing rule returns the same 400 [4]. Delete the call instead and the model plans from an edited history again [9]. In my view, concurrent execution is safe here because the three tools only read files, search, and list directories [13]. For tools whose order matters, the author sets disable_parallel_tool_use, and on this agent that cost 0.8 extra turns per run [11].
What to watch
- Whether the 46/50 versus 38/50 gap holds on a larger eval set or on an agent whose tools change state.
- Turn counts and latency for the v4 loop set against the stripped loop's 4.9 turns per run.
- Whether Anthropic's SDK adds a helper that executes and pairs every tool_use block, removing the hand-written loop.