Skip to content

Build1 publisher2 min readPublished

Eight of the 24 codes LLM vendors return with HTTP 429 are billing states only a human can clear

OpenAI and Qwen document eight billing and account errors under HTTP 429, a third of the 24 codes nine model vendors list there, a survey on dev.to found. Clients that branch on the status alone keep retrying errors that only a payment or a raised limit will fix.

The Engineer · Build desk

Illustration accompanying Eight of the 24 codes LLM vendors return with HTTP 429 are billing states only a human can clear

What happened

  • DeepSeek returns HTTP 402 when a balance runs out, while OpenAI sends a 429 for the same condition.
  • Qwen splits its purchase errors across statuses: an unpurchased workspace subscription is a 429, an unactivated Model Studio service a 403.
  • Anthropic's only 429 code, rate_limit_error, is documented to cover a rate limit, the usage tier's monthly spend cap, or a Claude Code workspace spend limit.
  • Google's own 402 remedy tells callers not to retry until credits are added; the author found no such line on the 429s that end up in retry loops.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost A status-keyed loop logs an unpaid bill as a transient capacity problem, so on-call engineers start from the wrong diagnosis and look at the vendor before the invoice.
  • exposure Multi-vendor gateways are the most exposed: one status-keyed handler treats the same unpaid bill as fatal at DeepSeek and retryable at OpenAI.
  • constraint A per-code table built from vendor docs still cannot tell an Anthropic rate limit from an Anthropic spend cap, so no static mapping cleanly separates retryable from fatal.

The loop has one input, the status. Catch it, sleep, double the sleep, try again. The author of the dev.to survey wrote that every SDK ships one and every tutorial recommends it [1]. Given a billing error, six attempts with exponential backoff take about a minute and then raise the error the first call already returned [6]. In the author's words, the loop "spends about a minute confirming that your credit card is still not on file" [6].

The billing label comes from the vendors themselves. The author took it from the remedy each vendor prints next to the code, such as OpenAI's "Add credits to continue using the API" for Credit balance exhausted [4]. Qwen's line for BudgetLimitExceeded is blunter: "You will be unable to make further API calls until the budget limit is increased or reset." [5]

The other 16 codes describe rates, quotas or capacity [2]. A few need a second look. Google's quota_exceeded is documented as a daily quota [13]. Counting it, I get nine of the 24 codes whose documented cause outlasts a one-minute backoff [3]. AWS's ModelNotReadyException says the model is not ready to serve inference requests. Waiting can plausibly fix that state [15]. Two Qwen codes, Throttling and Throttling.AllocationQuota, have no published cause, so any class assigned to them is a guess [14].

A third of the codes is a fact about documentation [1]. The share of one client's 429s that are billing states depends on its vendor mix and how often it hits a cap. The survey counted codes in published docs [2].

OpenAI and Qwen did one thing well. Their eight billing codes are distinct strings with printed remedies, so a lookup keyed on the vendor's code catches all of them [3][4]. Docs move, though. The author introduced Anthropic's documented cause for rate_limit_error as what it "now reads" [11].

The handler I would build runs in this order:

1. Parse the vendor's error code from the response body. The status only picks which parser to use. 2. Look the code up in a per-vendor table with three classes: transient, billing, ambiguous. 3. Retry transient codes with backoff. 4. Fail fast on billing codes and page a human, with the vendor's remedy text in the alert. 5. Give ambiguous codes, Anthropic's rate_limit_error among them, a short retry budget, then escalate them as billing [12].

What to watch

  • Whether Anthropic moves the spend-cap and workspace-limit cases out of rate_limit_error into their own code or status.
  • Whether OpenAI or Qwen add a "Don't retry" line to their billing 429s, matching the wording Google prints on its 402.
  • Whether SDKs that retry on 429 by default start excluding the documented billing codes.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories