Build1 publisher3 min readPublished
The refund button is the architecture: inside the tool-use layer of a support agent
A published production loop for customer service agents shows where the cost really sits: not the model, but the classify-execute-confirm middle where a write hits a payment processor.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Most customer service AI implementations retrieve relevant information from a knowledge base, synthesize a response, and hand the conversation back to the customer; the post calls that a chatbot.
- A customer service agent that executes workflows processes the refund, updates the account, triggers the return label, and sends confirmation when it is done.
- The post states the difference is not the model but the architecture around the model, specifically the tool-use layer and how everything connecting to it is designed.
- A Q&A agent has one primary operation: retrieve context, generate response.
- A workflow-execution agent has three operations: classify intent, execute tools, generate response; the middle step is where production complexity lives.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Dextra Labs has published, on dev.to, the agent loop it says it uses for production customer service agents, framed as the part most tutorials skip [15]. The useful content is the boundary it draws: a retrieval agent synthesises an answer and hands the conversation back, while an execution agent processes the refund, updates the account, triggers the return label, and confirms when it is done [1][2].
The post's own summary of the difference is that it is not the model but the architecture around it, specifically the tool-use layer [3]. That matches the operation count. A question-answering agent has one primary operation, retrieve then generate [4]. A workflow agent has three: classify intent, execute tools, generate response, and the middle one is where production complexity lives [5].
The loop as published runs six steps: classify intent, retrieve customer context, plan actions against a loaded policy, execute tools, generate a grounded response, persist state [6]. Step four is the one that executes real operations against real systems [17]. Escalation is checked after tool execution and before response generation, with the conversation, context, tool results and a reason handed to the human [7].
Classification is a separate model call: claude-sonnet-4-5, capped at 256 tokens, last three turns of history, returning JSON with a type, a confidence and entities across eight categories [9][8]. So the minimum cost of a turn is two model calls before any tool runs [16]. The design decision worth noting is what happens to that confidence number. Below 0.70, the post says, tool permissions are scoped tighter and the escalation threshold drops [10]. That turns a self-reported score into an authorisation input, which is a defensible pattern and also the sort of number that needs calibration data behind it. The excerpt states the threshold without showing how it was chosen [18].
The tool layer is described as the interface to CRM, order management, payment processor and helpdesk [11]. Two definitions are shown. lookup_order takes an order id and a customer id and returns status, tracking and line items [12]. process_refund is described as initiating a refund for an eligible order within policy parameters, taking order id, refund amount and reason [13]. One reads, one moves money [19]. The published listing cuts off partway through the refund tool's schema [14], which means the interesting half of the write path, the part that decides whether a retry issues a second refund, is not on the page.
Two things to watch in this pattern, whoever is selling it. First, where the policy lives. In the loop as listed, load_policy is called during planning, at step three, before execution at step four [20]. That puts the eligibility constraint upstream of the tool rather than inside it, so the backend is trusting a plan rather than enforcing a limit. Second, the ordering of escalation. Checking escalation conditions after tool results are in [7] means the human arrives after the refund has already been attempted, which is the right sequence for reporting and the wrong one for prevention. The confirmation step that the post treats as the finish line [2] is the cheapest part of the build.