Build1 publisher3 min readPublished
Deferred tool schemas cut cost 21% on average, and made one task type 12.3% dearer
Progressive disclosure on a 20-tool agent saved about 30% of input tokens overall, but the per-task split shows cost stops being flat and starts tracking your traffic mix.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- A team turned on progressive disclosure for an agent with 20 tools; deferring the tool schemas cut input tokens by about 30% and total cost by about 21%.
- Two arms were compared: always-on, with all 20 schemas serialised into every request, and deferred, with all 20 wrapped in Pydantic AI's DeferredLoadingToolset so the model has to search for a capability and load it before calling it.
- Three of the four task types saved 21% to 30%; the fourth saved nothing on one transport and cost 12.3% more on the other.
- All 80 runs are public and the experiment can be reproduced offline for free.
- The benchmark used four business tasks, each needing exactly one tool, with twenty tool schemas registered always.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A team switched on progressive disclosure for an agent with 20 registered tools, deferring the schemas until the model searches for a capability, and measured input tokens down about 30% and total cost down about 21% [1]. Then, according to the dev.to writeup, they split the same 80 runs by task type and the average came apart: three of the four tasks saved between 21% and 30%, while the fourth saved nothing on one transport and cost 12.3% more on the other [3].
The setup is small enough to argue with. Four business tasks, each needing exactly one tool, with all 20 schemas registered in both arms [5]. The always-on arm serialised every schema into every request; the deferred arm wrapped the same 20 in Pydantic AI's DeferredLoadingToolset, so the model had to search for a capability and load it before calling [2]. Twenty runs per cell, sequential, no retries, gpt-4o at temperature 0 with parallel_tool_calls disabled, run on 2026-08-06 against pydantic-ai-slim[openai] 2.24.0 [6]. Both Chat Completions and the Responses API were tested, because tool search executes server-side on one and through a local fallback on the other [7]. All 80 runs are published, and the whole experiment cost $0.279585, about $0.0035 a run [4][8][5].
Pooled, the result looks clean. Input tokens fell 30.6% on Chat Completions and 26.1% on Responses, cost fell 21.2% and 16.8%, and correctness held at 80 of 80 exact matches [10][11]. The tell is the standard deviation column: 10.3 and 10.6 tokens in the always-on cells against 190.8 and 263.1 in the deferred cells [9], which is 19 to 25 times more variable [4].
The author's first assumption was run-to-run nondeterminism, and reports that it was wrong [16]. Grouped by task inside each cell on the Responses API, always-on puts all four tasks within 22 tokens of each other, every one with a standard deviation of 0.0, because the 20 schemas dominate the prompt [13]. Under deferral the same four tasks spread from 807.0 to 1,431.8 mean input tokens, a gap of 625 tokens or 1.77x between cheapest and dearest [14][3]. Inventory-reorder fell 40.0% on input tokens, shipment-delay 38.8%, currency-conversion 30.7% [2]. Defect-threshold rose 5.0% [1], and it is also the only task with meaningful run-to-run spread under deferral, standard deviation 95.6 against 0.0 or 2.7 for the others [14]. Retrieval that is unstable and retrieval that is expensive appear to be the same failure here.
The reframing is the part worth keeping: deferral does not make cost noisy, it makes cost task-dependent, so the bill starts tracking traffic mix rather than tool count [15]. Always-on charges the same amount whatever the user asks [13].
So watch the mix before you quote a savings figure. If defect-threshold-shaped traffic is a tenth of your volume the mean survives; if it is most of your volume it does not, since that task type went backwards rather than flat [3]. Watch the transport separately, because the 12.3% regression showed up on one and a wash on the other [3], which is consistent with tool search running server-side in one path and locally in the other [7]. And watch the sample size: five runs per task per cell [13], one model, one library version [6], one 20-tool agent [5]. The numbers may not transfer. The method - group by task and by transport before believing the average - costs about $0.28 to repeat [8].