Build1 publisher3 min readPublished
Retry Loops Fail Because They Classify Nothing: A Per-Failure-Class Taxonomy
A dev.to post argues four different failures arrive as one event in most scrapers. The useful part is not the backoff curve but the counter you enforce on yourself before the server does.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Most scrapers have a single retry policy: catch the exception, sleep, try again, give up after N attempts. It is the default in every tutorial, tenacity provides it in three lines, and it is fine on a healthy target. The post's argument is that this policy makes blocks worse.
- A rate limit, a hard block, a TLS reset and a slow origin all arrive as "the request failed"; treating them identically means spending the entire run budget hammering something that was never going to open, while the one failure that would have succeeded on retry gets three attempts and then gets dropped.
- The author states that in this taxonomy "the counter matters more than the backoff".
- For a 429 with a Retry-After header: honour it, sleep the stated duration, then continue on the same connection and the same identity. Do not rotate anything, because rotating there is how a temporary throttle turns into a fingerprinted pattern.
- For a 429 with no header, exponential backoff with jitter is the right tool, and the author advises starting higher than expected: 5 seconds, not 0.5.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A post on dev.to argues that the retry loop in most scrapers is not merely wasteful but is part of what earns the block [1]. The claim is worth an operator's attention because it moves the important decision away from sleep duration and towards classification: a rate limit, a hard block, a TLS reset and a slow origin all arrive as "the request failed", so you burn the run budget on a target that was never going to open while the one failure that would have succeeded gets three attempts and is dropped [2].
The taxonomy has four entries. A 429 carrying a Retry-After header is the cooperative case: sleep the stated duration and continue on the same connection and the same identity, because rotating there is how a temporary throttle becomes a fingerprinted pattern [4]. A 429 with no header is the one place exponential backoff with jitter is genuinely right, and the author says to open at 5 seconds rather than 0.5 [5]. A 403, a 401 or a challenge page is not retryable at all: the same bytes get the same rejection, and the repeat is what converts a temporary block into a durable one, so the move is straight to a fallback identity, a different route, or shelving the URL [6]. Connection resets, timeouts and TLS handshake failures are the class people wrongly file with 403; they are usually noise, worth two immediate retries and then a hard failure [7]. The dividing line the author draws: HTTP rejections tell you something about your request, transport failures usually tell you nothing [8].
The published code is tighter than the prose in one place and looser in two. Its hard-block set includes 407 and 451, neither of which the written taxonomy discusses [4]. Its Retry-After path only fires when the header value passes `isdigit()`, so a date-formatted Retry-After silently falls through to the exponential ladder and you ignore the number the server gave you [5]. That ladder is `min(60, 5 * 2 ** attempt)` plus jitter, which means the 60-second cap engages at attempt four, while the 5xx ladder reaches its 30-second cap at attempt five [10][3]. The transport path sleeps 0.5 seconds and gets two tries, so the whole class costs about one second before you give up on it, against five seconds for the opening sleep on a headerless 429 [2].
The load-bearing argument is the next one. Backoff handles the failure you already have and does nothing about the next one, and on many targets the limit is not time-based at all but a volume budget: N requests, block on N+1, however politely you spaced them [11]. The author's example is a regional grocery chain's stock checker that would 429 politely for about an hour and then flip to a permanent 403 once some invisible total was crossed, which no spacing strategy touches because the thing being counted is volume, not rate [12]. The proposed fix is a rolling per-host counter that refuses before the server does, defaulting to 200 requests in a 3600-second window, or roughly one request every 18 seconds [13][1].
Once that counter exists, a 403 stops being something you react to and becomes a number you write to disk: whatever the counter read when the block landed is the new ceiling for that host, minus margin, and the next run starts already knowing [14]. The author's reason for distrusting the 403 on its own is timing. By the time it arrives you are already benched, the cooldown clock started without telling you, and nothing you do for the next hour counts [15].
Two things to check in your own stack. Whether your classifier honours non-numeric Retry-After values, and whether any host ceiling you have ever hit was recorded anywhere durable rather than lost with the process.