Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished

Recurring invalid_grant is usually two of your own workers racing one refresh token

A dev.to writeup argues the intermittent OAuth failure is a concurrency bug on a rotating token, and that retrying it looks exactly like the replay attack reuse detection exists to catch.

The Engineer · Build desk

How we use AISend a correction

What happened

  • A dev.to writeup argues that intermittent OAuth invalid_grant failures usually come from two of your own processes redeeming the same refresh token at once.
  • Under rotation, the losing call presents an already-consumed token, and many providers read that as replay and revoke the entire token family.
  • The spec crams invalid, expired, revoked, redirect-URI-mismatched and wrong-client grants into the single invalid_grant string.
  • Rotation is recommended in the OAuth 2.0 Security Best Current Practice and is the default in OAuth 2.1 drafts.
  • The recommended remedy is single-flight refresh per integration plus shared token state, rather than each worker noticing expiry on its own.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Retry budgets are useless here and actively harmful: five attempts produce five replays, so the mitigation you already own is the one you must switch off.
  • exposure The bill lands on the customer and your support queue, because a family revocation ends with a person re-authorising an account that was healthy a minute earlier.
  • decision Provider documentation stops being the input to the design; you decide whether to serialise based on what a live token response returns for your specific app registration.
  • capability Teams already on Postgres can serialise refresh inside the transaction they have, without standing up a lock service or writing lease-expiry logic to get wrong later.

Triage order matters more than the taxonomy. Of the four things `invalid_grant` can mean on a refresh call, three are steady states: an idle-expired refresh token, a revoked grant, and client credentials that do not match the token will fail every time you try them [6]. Only reuse-after-rotation comes and goes [7]. So the useful first question is not what the string means, it is whether the errors are bursty, whether they cluster on traffic spikes or cron minute boundaries, and whether they hit accounts that were working ten minutes ago [7].

The arithmetic explains why this survives staging. One process locally; a web dyno, three queue workers and a scheduled job in production, all reading the same row [10], which is five candidates for the same refresh [18]. The refresh path opens two minutes early, because the fast path only hands back the cached access token while `expires_at` minus a 120-second skew is still in the future [3]. That gives a 120-second band in which any of those five will independently decide the token is stale [15]. The overlap in the worked example is 40 ms wide [2], roughly three hundredths of a percent of the band [16]. A single-process dev box never rolls that dice.

Retrying is the wrong layer, and the writeup is blunt about the reason: a retry loop turns one replayed token into five, which is indistinguishable from the attack reuse detection was built to catch [11]. The damage is also not confined to the loser. Where the provider revokes the family, worker A's brand-new token dies alongside B's replayed one [2], and the recovery path runs through a human pressing Reconnect [20].

The lock itself is the cheap part. An advisory transaction lock keyed on a stable hash of the integration id needs no new infrastructure, has no lease expiry to reason about, and releases when the transaction ends, including when the connection dies [13][4]. The step that does the actual single-flighting is the re-read inside the lock [4]; serialising the workers without re-checking state would queue five refreshes rather than running them concurrently, which is the same number of redemptions arriving in a politer order. The configured 15-second `lock_timeout` is the ceiling on how long a queued worker waits before the statement gives up [4][17].

One caution about the evidence. This is a single writeup with no incident data attached, and it says many providers treat reuse as replay without naming them [20]. It also says the behaviour can differ per app registration, which makes provider documentation the weaker authority [8]. The check is one refresh call: compare the `refresh_token` in the response against the one you sent, and if it came back changed, every refresh from then on is a state mutation you have to serialise [9]. If it came back identical, you are chasing one of the three steady causes instead [6].

What to watch

  • Whether providers begin documenting what reuse detection actually revokes, the single token or the whole family, instead of leaving integrators to infer it from a dead account.
  • Whether OAuth 2.1 shipping rotation as the default flips providers you currently treat as non-rotating, widening the set of integrations that need serialised refresh.
  • Whether advisory transaction locks hold their auto-release guarantee behind a transaction-mode connection pooler, where the transaction boundary is not the application's.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence42
Adoption
Insufficient
Hype gap+22
Incentives
Insufficient
Confidence45
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    The prescribed fix is to make refresh a single-flight operation per integration and to stop treating access token expiry as something every worker discovers independently.

  2. [2]

    In the worked race, worker A reads the row, sees expires_at in the past and POSTs; worker B reads the same row 40 ms later, sees the same stale expires_at and POSTs the same refresh token; A gets a new pair, B presents an already-redeemed token and gets invalid_grant, and depending on the provider the family is revoked and A's brand-new token is dead too.

  3. [3]

    The sample code sets REFRESH_SKEW to 120 seconds and takes a fast path that returns the stored access token with no lock and no network call when expires_at minus the skew is still in the future.

Sources

1 independent publisher whose own reporting we read for this story.

  1. dev.to

    1 article · August 24, 2026

    Why Your OAuth Integration Randomly Returns invalid_grant (and How to Stop Two Workers From Racing)

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

  • OAuth refresh-token rotationFollow
  • Postgres advisory locksFollow
  • Third-party API integration reliabilityFollow
  • Distributed concurrency controlFollow
Loading related stories