Build1 publisher3 min readPublished
Anthropic streams tool arguments as JSON fragments, so pick a coping strategy on purpose
A dev.to field-notes post lays out how partial_json actually arrives, and why the three common ways of handling it fail in three different places.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- With the Anthropic API, a streamed tool call does not show up as one JSON blob; the arguments arrive as raw JSON text fragments across events.
- The material is a field-notes post published on dev.to under the headline "Streaming tool calls without losing your mind", originally published on Loop & Retry, described as field notes on building LLM agents that survive production.
- A streamed Anthropic tool call appears as a content_block_start of type tool_use (with a name and an empty input), followed by a run of content_block_delta events whose delta.type is input_json_delta, each carrying a fragment of the arguments as raw text in partial_json.
- "You get the characters of the JSON object, not the object."
- After three deltas the accumulated buffer might read {"target": "sta ; feeding that to json.loads produces a JSONDecodeError, every time, until the block actually closes.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A set of field notes published on dev.to under the headline "Streaming tool calls without losing your mind" draws a line that most agent codebases cross on the first pass: with the Anthropic API, a streamed tool call does not arrive as one JSON blob, it arrives as raw JSON text fragments [2][1]. Anything that renders or acts on the accumulated buffer before the block closes is operating on invalid JSON, and per the post it fails every time rather than occasionally [5].
The event shape is specific. A `content_block_start` of type `tool_use` carries a name and an empty `input`, then a run of `content_block_delta` events whose `delta.type` is `input_json_delta` each carry a fragment of the arguments as raw text in `partial_json` [3]. As the author puts it, you get the characters of the JSON object, not the object [4]. Three deltas in, your buffer might read `{"target": "sta`, and `json.loads` on that raises `JSONDecodeError` until the block actually closes [5]. Only at `content_block_stop` is the accumulated string guaranteed parseable [6]. That is a parser behaving correctly on invalid input, not a bug in your handling [7].
The reason teams walk into this is that plain-text streaming is a solved problem, because a partial sentence is still readable, and tool calls look like the same job [8]. They are not, and the post enumerates three coping strategies that trade off differently [9]. Two of the three it rejects outright, and the third it accepts only for execution [2].
First: accumulate, call `json.loads` after every chunk, catch the exception, continue. It does not crash, but you are running exception-driven control flow on the hot path of every tool call, dozens of times per call, and it hides the one exception that matters, a genuinely malformed final payload; the log line you care about becomes indistinguishable from noise [10].
Second: buffer everything and parse once at `content_block_stop`. Correct, and the right thing for execution, but if that is all you do you have opted out of streaming for tool calls while your text responses still stream token by token [11]. On a tool call with a large argument, a long file body or a multi-paragraph message draft, the user watches nothing for the entire generation and then sees the whole result at once [12].
Third: guess the shape by string matching, tracking open braces, counting quotes, treating a comma at depth 1 as the end of a value. It survives the happy path and breaks on the first argument value containing a brace, an escaped quote, or a comma of its own, which for free text is a question of when [13].
The recommended split is to stop treating parse-for-display and parse-for-execution as one operation, because they have different tolerance for being wrong [14]. Display wants a tolerant read of an incomplete document, good enough for a progress skeleton and never good enough to act on [15]. The post's `best_effort_partial` helper appends a closing quote when the quote count is odd, closes whatever brackets remain open, and returns `None` when the result still will not parse, with a docstring warning never to feed the output to anything that executes [16]. Run it per delta, render an incrementally filling form, and skip the frame when it returns `None` [17]. That leaves two parse paths per tool call with two contracts [1].
Worth checking in your own stack: whether a swallowed `JSONDecodeError` sits on the tool-call hot path [10], whether any display-side tolerant parse can reach an executor [16], and which of your tools carry arguments large enough that end-of-block parsing shows up as dead air [12].