Build1 publisher2 min readPublished
Hard spend caps on AWS and Google Cloud trade a runaway bill for an outage
AWS added a per-project spend limit on September 16, 2026 that pauses service at the monthly cap, following Google Cloud's July launch of Spend Caps. A cap that checks each request stops at once, while one built on lagging billing data keeps charging until a function fires.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Simon Willison published an essay on October 3, 2026 arguing that almost no pay-per-use service stops spending on its own when usage runs away.
- A classic billing alert is an email or webhook that warns of a crossed threshold and interrupts nothing, and it is still the default in several cloud consoles.
- Under a hard cap, the provider rejects every new request until the account owner lifts the limit.
- OpenAI and Anthropic already let account owners set a hard spend cap from each account's billing panel.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- exposure A customer-facing project under a cap can now be paused by its own billing limit, so a cost overrun turns into an availability incident with a billing trigger.
- cost A per-request cap can still overshoot: the counter is checked before each call but updated only as responses finish, so calls already admitted add spend the check never saw.
- contradiction The post calls Google's Spend Caps a per-service cap inside a project, yet says cloud caps meter all of an account's services, and that scope decides what a tripped cap takes down.
A hard cap can be built two ways, and the choice sets how much a runaway costs before it stops. The first keeps a usage counter and checks it on each request, before the request is processed [8]. The cut happens inside that same request, so it is close to instant [8]. The dev.to post says OpenAI and Anthropic work this way, and cites Willison for the view that Google Cloud's new Spend Caps do too [8].
The second design is the classic AWS or Google Cloud budget. It reads billing data that is consolidated with a delay and checks accumulated spend at intervals [9]. When spend crosses the threshold, it can trigger a Lambda or a Cloud Function that revokes access or disables billing, but only if someone configured that [9]. This path is never instant. Spend keeps accumulating between the moment the threshold is crossed and the moment the function runs [9]. The post does not say which design AWS's September 16 limit uses, and I would want that settled before counting on it to stop a loop quickly.
The meters differ as well. OpenAI and Anthropic count estimated dollars, computed from input and output tokens and updated as each response finishes generating [10]. AWS and Google Cloud count real accumulated spend across every service in the account [11]. A limit there can cut compute, storage and network at the same time, along with the model calls [11].
Willison's argument is about the default [2]. He wants the hard cut switched on for every pay-per-use service, with an explicit checkbox for anyone who prefers to keep running past the limit and accept the later charges [12]. In the post's Spanish rendering, he frames the choice as errors in the middle of the night or a bill of several thousand dollars the next morning [14].
The post's worry is coding agents. One wired to a real cloud account can misread an instruction and start creating instances or calling an expensive model in a loop [13]. Without a hard cap, nobody finds out until the month-end billing summary arrives [13].
I think any project an agent can reach with real credentials should carry a hard cap, and so should sandboxes and test accounts. An outage there costs a rerun. The post describes AWS's limit as working per project [3]. Under that design, a team keeps agent experiments in their own project, apart from anything customers depend on.
What to watch
- Whether AWS or Google Cloud switches the cap on by default for new projects with an explicit opt-out, the change Willison asked for.
- What a paused AWS project retains while its cap is in force, such as stored data and running resources.