Skip to content

Build1 publisher3 min readPublished

Nobody chose retry-by-default, and the bill arrives as your customer's timeout

A field-notes post argues the retry question splits in two: how many, and whether at all. On a user-facing path with a 30-second ceiling, the honest answer is usually a fast error.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • The post distinguishes a prior retry-budget piece, which answered HOW (a shared bucket that caps retries as a fraction of throughput), from this one, which answers WHEN a call should fail immediately instead of burning that budget; the answer depends on what fails, who is waiting, and what happens next.
  • Most code does not ask when to give up; it retries by default, on some combination of "it might work next time" and "the library gives me retry for free". The consequence is that retry-ready code becomes the customer's timeout, which becomes the platform's cascade.
  • A request from a user interface has a hard deadline: a browser will not wait longer than about 30 seconds before the user is left with a spinner and closes the tab.
  • For user-facing operations the post's cost model is: cost of one retry attempt is approximately seconds_added_to_response_time x 1; cost of failing now is approximately user_abandonment_rate x minutes_of_lost_engagement x revenue_per_engagement.
  • Amazon measured its own checkout and found each 100ms of added latency costs roughly 0.1% of sales.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A post published on dev.to under the Loop & Retry banner separates two questions most codebases collapse into one: how much retrying you can afford, which it treats as a shared budget capping retries as a fraction of throughput, and whether a given call should retry at all [1]. Its useful observation is that the second question is rarely answered on purpose, because retry arrives free with the client library and stays [2].

The user-facing case is the one where the default is most clearly wrong. A request from a user interface has a hard deadline: roughly 30 seconds before the browser is just a spinner and the tab closes [3]. Inside that window you are choosing between a fast error and a long hang that ends in the same error, and the post frames it as a cost comparison: seconds added to response time on one side, abandonment rate times lost engagement times revenue per engagement on the other [4]. According to the author, Amazon measured its own checkout and found each 100ms of added latency cost roughly 0.1 percent of sales [5]. Five seconds of retry is fifty of those increments, which arithmetic puts at about 5 percent of sales [6][7] - a larger number than the odds that a second attempt beats an error already firing on the first [8]. Hence the rule: fail fast unless the error is known-transient and rare, such as a network timeout or a temporary 503 [8]. Most services, the post says, land on a 30-second ceiling with two quick retries at 1s and 3s backoff [9], which spends four seconds of waiting, about 13 percent of the ceiling [10].

The placement argument matters more than the numbers. Retry belongs at the request handler, not in library code, because if the client retries down in userspace the handler sees the third attempt as a fresh call and has no decision left to make [11][12].

Background jobs invert the tradeoff, and for a reason worth internalising: a call-level retry keeps partial progress, while giving up and re-queueing makes the supervisor re-run the whole task [13][14]. On a ten-minute job that hits a transient error at minute eight, a five-second retry costs five seconds and a re-queue costs ten minutes plus overhead [15], roughly 120 times more [16]. So retry longer, with a ceiling: if the service is actually down, a hundred retries pile up behind each other and block the rest of the queue [17]. The recommended shape is short backoff, 1s, 2s, 4s, to a total of 30 to 60 seconds, then escalate to a human alert and wait on it with never-ending exponential backoff rather than burning the queue with re-queues [18]. Again the placement: at the call site inside the job, since job-level retry is for code bugs and bad input, not transient service errors [19].

The reason any of this is load-bearing is amplification. N concurrent callers each retrying on their own turns a hiccup into a flood, and the post's example has a service built for 100 RPS seeing 300 [20] - two extra requests for every real one, arriving exactly when the dependency is weakest [21].

What to watch in your own stack: whether your HTTP client is retrying below the layer that owns the deadline, and whether you can measure the ratio of attempts to requests during an incident rather than guessing it. The Amazon latency figure here is the author's citation, not something we verified, and the source text cuts off mid-cascade, so treat the amplification section as a sketch rather than a measurement.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories