Build2 publishers3 min readPublished Updated
Agent budget gates must reserve units before the paid tool call runs
Agent runtimes should refuse unit 501 of a 500-unit budget before the call runs, a dev.to post argues. The new AWS and Google Cloud spend caps then become the backstop behind a tighter limit in the agent's own call path.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Simon Willison wrote on October 3 that usage-priced services need hard caps that cut off and return errors by default, with removing the cap an explicit opt-in.
- AWS announced on 16 September that when a project's usage reaches its monthly spend limit, the project is paused for that month.
- Google Cloud's Spend Caps preview can pause eligible service usage once a configured cap is enforced, while the project's resources stay intact.
- The dev.to post's sample checks each reservation against a budget object and raises an error before the tool runs if the limit would be exceeded.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost An AWS project that hits its limit stays paused for the rest of the month, so one looping agent halts every other workload that shares the project.
- exposure Existing AWS accounts outside the limited rollout have no AWS hard cap yet, so for them the application's own counter is the only hard stop.
- precedent If coding agents start steering builders toward providers with hard caps, as Willison hopes, cap support becomes a selection criterion for agent-built deployments.
The failure the post describes looks like ordinary work. Each action an agent takes can be permitted while the run as a whole becomes economically unsafe [1]. In the post's example, one bad decision becomes twenty valid paid actions, and a second recovery rule turns those twenty into a hundred [2]. That is 100 paid actions from one choice, and the second rule alone multiplies the count by five [1].
An alert watches that loop and lets it keep going. The post writes the difference as two pipelines. A soft budget runs observe, alert, keep running. A hard budget runs observe, compare, deny the next action [3]. Willison put the buyer's side plainly. "I expect that most businesses and individuals would prefer errors to a surprise $10,000+ bill," he wrote [11].
Provider caps are the coarse layer. Google's Spend Caps launched in July and set a monthly cap on specific services within a project [15]. According to the post, Google warns that enforcement is not instantaneous because it works from estimated costs [6]. The post does not say how long that delay runs. Whatever an agent spends between the estimate and the pause is already on the bill. For that reason the post wants a tighter limit held by the application itself [6].
In the sample code, the gate is one comparison inside `reserve()`: `if self.used + units > self.limit: raise RuntimeError("hard budget exceeded")` [7]. Only after that line passes does `call_metered_tool` invoke the tool [7]. According to the post, a check placed after the call is only telemetry, a retry path that skips the wrapper is not a hard cap, and a child agent given a fresh counter can escape the limit by spawning [8].
The post calls the sample intentionally small [7], and two gaps follow from that. It reserves `estimated_units` and never settles the reservation against what the call actually cost. A call that runs over its estimate therefore pushes real usage past the limit [3]. The application inherits Google's estimation problem. The check-then-increment also runs without a lock, so two concurrent child agents sharing one `Budget` could both pass the comparison before either one increments [4]. In my view both gaps are cheap to close: an atomic counter, and a settle step after each call that charges the actual cost.
Money is one dimension of several. The post's example run budget allows 80 tool calls, 12 browser writes, 500 external API units, 1,200 seconds of wall clock and three child agents [9]. The wall-clock ceiling is 20 minutes [2]. A provider cap counts dollars per service [15]. A browser write or a spawned child exists only inside the runtime. I think the post has the layering right for any agent that can create its own work. It places the budget object at the same architectural level as authentication, authorization and cancellation [16].
What to watch
- Whether AWS moves its project spend limit from a limited rollout to general availability for existing accounts.
- Whether Google publishes how far Spend Caps enforcement can trail actual spend as the feature leaves preview.
- Whether agent frameworks ship a parent-scoped budget object on the tool-call path as a default.