Skip to content

Build1 publisher3 min readPublished

The refund that fired three times: tool calls are a systems problem, not a prompt problem

A developer's LLM issued the same refund three times after a timeout. The defect was in the wiring, not the model, and the fixes are schema validation and scoped toolsets.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying The refund that fired three times: tool calls are a systems problem, not a prompt problem
Generated illustration

What happened

  • The author of the dev.to article "Letting an LLM call your APIs without losing sleep" says the first time he gave a language model access to a real API it worked perfectly in a demo: it looked up an order, summarized the status, and everyone in the meeting nodded.
  • Two weeks later, in production, the same setup tried to issue the same refund three times in a row because a timeout made it think the first attempt had failed.
  • The author writes that nothing was "wrong" with the model, and that everything was wrong with how he had wired it up.
  • Function calling (also called tool calling) is described as the moment an LLM stops being a text generator and becomes an actor in your system.
  • When a model "calls a function" it emits text that looks like JSON; the SDK parses it and hands you an object, and that object has the epistemic status of user input from a public form.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A developer writing on dev.to handed a language model access to a real API, demoed it to a nodding room, and two weeks later watched the same setup try to issue the same refund three times in a row because a timeout made it think the first attempt had failed [1][2]. That is the case for treating tool calling as a systems problem: the moment you pass a function schema, the model stops being a text generator and becomes an actor inside your system [4].

The author is blunt that the model was not the defect. Nothing was wrong with the model; everything was wrong with how he had wired it up [3]. The triple refund is a retry story. Most agent frameworks and most hand-rolled loops have retry logic somewhere, and a thrown exception looks like transient infrastructure failure, so something retries it: the framework, the queue, or your own catch block [17]. He says the incident cost him real money [18].

Worth being clear about the limits of the source: the published text stops mid-sentence as it starts walking through a payment API timing out, so it never states the remedy [20]. What survives is the shape of the problem. Any deduplication has to sit on the same side of the boundary as the payment call, because the retry can originate from three different layers, none of which knows the first request landed [17].

The parts the piece does finish are about narrowing what the model can reach. A function call is text that looks like JSON, parsed by your SDK into an object, and that object has the epistemic status of user input from a public form [5]. The failure mode is arguments that are almost right: a string where a number belongs, an ISO date in the wrong timezone, an enum synonym such as "cancelled" when your API says "canceled", a negative quantity, an ID copied from the wrong part of the conversation [6]. His rule is schema-first, validated at runtime, with one schema deriving both the tool definition sent to the model and the validator that guards execution [7]. Business limits go in the schema, not the prompt: his example caps amountCents at 50,000, which makes a single refund structurally incapable of exceeding 500 dollars [8][9][19]. Prompts are suggestions; validators are laws [9].

The second half is capability scoping, which is the cheaper control. Read and write tools are different risk classes, and a conversation that only answers questions gets only read tools, because a model cannot misuse a tool it was never given [12]. Account identifiers are injected server-side; if the model can pass a customerId, it can pass the wrong one [13]. The refund tool enters the toolset only after an order has been located and verified, which shrinks the blast radius of every earlier turn [14]. He argues this also improves accuracy, since tool selection is a decision the model can get wrong and every irrelevant tool is another chance to get it wrong [15].

Watch whether your agent framework retries tool executions by default, and at which layer. Then check whether your write endpoints can tell a repeat call from a new one.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories