Skip to content

Build1 publisher2 min readPublished

Hard spend caps on AWS and Google Cloud trade a runaway bill for an outage

AWS added a per-project spend limit on September 16, 2026 that pauses service at the monthly cap, following Google Cloud's July launch of Spend Caps. A cap that checks each request stops at once, while one built on lagging billing data keeps charging until a function fires.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Hard spend caps on AWS and Google Cloud trade a runaway bill for an outage
Generated illustration

What happened

  • Simon Willison published an essay on October 3, 2026 arguing that almost no pay-per-use service stops spending on its own when usage runs away.
  • A classic billing alert is an email or webhook that warns of a crossed threshold and interrupts nothing, and it is still the default in several cloud consoles.
  • Under a hard cap, the provider rejects every new request until the account owner lifts the limit.
  • OpenAI and Anthropic already let account owners set a hard spend cap from each account's billing panel.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure A customer-facing project under a cap can now be paused by its own billing limit, so a cost overrun turns into an availability incident with a billing trigger.
  • cost A per-request cap can still overshoot: the counter is checked before each call but updated only as responses finish, so calls already admitted add spend the check never saw.
  • contradiction The post calls Google's Spend Caps a per-service cap inside a project, yet says cloud caps meter all of an account's services, and that scope decides what a tripped cap takes down.

A hard cap can be built two ways, and the choice sets how much a runaway costs before it stops. The first keeps a usage counter and checks it on each request, before the request is processed [8]. The cut happens inside that same request, so it is close to instant [8]. The dev.to post says OpenAI and Anthropic work this way, and cites Willison for the view that Google Cloud's new Spend Caps do too [8].

The second design is the classic AWS or Google Cloud budget. It reads billing data that is consolidated with a delay and checks accumulated spend at intervals [9]. When spend crosses the threshold, it can trigger a Lambda or a Cloud Function that revokes access or disables billing, but only if someone configured that [9]. This path is never instant. Spend keeps accumulating between the moment the threshold is crossed and the moment the function runs [9]. The post does not say which design AWS's September 16 limit uses, and I would want that settled before counting on it to stop a loop quickly.

The meters differ as well. OpenAI and Anthropic count estimated dollars, computed from input and output tokens and updated as each response finishes generating [10]. AWS and Google Cloud count real accumulated spend across every service in the account [11]. A limit there can cut compute, storage and network at the same time, along with the model calls [11].

Willison's argument is about the default [2]. He wants the hard cut switched on for every pay-per-use service, with an explicit checkbox for anyone who prefers to keep running past the limit and accept the later charges [12]. In the post's Spanish rendering, he frames the choice as errors in the middle of the night or a bill of several thousand dollars the next morning [14].

The post's worry is coding agents. One wired to a real cloud account can misread an instruction and start creating instances or calling an expensive model in a loop [13]. Without a hard cap, nobody finds out until the month-end billing summary arrives [13].

I think any project an agent can reach with real credentials should carry a hard cap, and so should sandboxes and test accounts. An outage there costs a rerun. The post describes AWS's limit as working per project [3]. Under that design, a team keeps agent experiments in their own project, apart from anything customers depend on.

What to watch

  • Whether AWS or Google Cloud switches the cap on by default for new projects with an explicit opt-out, the change Willison asked for.
  • What a paused AWS project retains while its cap is in force, such as stored data and running resources.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories