BuildNot yet confirmed elsewhere1 publisher3 min readPublished
Recurring invalid_grant is usually two of your own workers racing one refresh token
A dev.to writeup argues the intermittent OAuth failure is a concurrency bug on a rotating token, and that retrying it looks exactly like the replay attack reuse detection exists to catch.
The Engineer · Build desk
What happened
- A dev.to writeup argues that intermittent OAuth invalid_grant failures usually come from two of your own processes redeeming the same refresh token at once.
- Under rotation, the losing call presents an already-consumed token, and many providers read that as replay and revoke the entire token family.
- The spec crams invalid, expired, revoked, redirect-URI-mismatched and wrong-client grants into the single invalid_grant string.
- Rotation is recommended in the OAuth 2.0 Security Best Current Practice and is the default in OAuth 2.1 drafts.
- The recommended remedy is single-flight refresh per integration plus shared token state, rather than each worker noticing expiry on its own.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint Retry budgets are useless here and actively harmful: five attempts produce five replays, so the mitigation you already own is the one you must switch off.
- exposure The bill lands on the customer and your support queue, because a family revocation ends with a person re-authorising an account that was healthy a minute earlier.
- decision Provider documentation stops being the input to the design; you decide whether to serialise based on what a live token response returns for your specific app registration.
- capability Teams already on Postgres can serialise refresh inside the transaction they have, without standing up a lock service or writing lease-expiry logic to get wrong later.
Triage order matters more than the taxonomy. Of the four things `invalid_grant` can mean on a refresh call, three are steady states: an idle-expired refresh token, a revoked grant, and client credentials that do not match the token will fail every time you try them [6]. Only reuse-after-rotation comes and goes [7]. So the useful first question is not what the string means, it is whether the errors are bursty, whether they cluster on traffic spikes or cron minute boundaries, and whether they hit accounts that were working ten minutes ago [7].
The arithmetic explains why this survives staging. One process locally; a web dyno, three queue workers and a scheduled job in production, all reading the same row [10], which is five candidates for the same refresh [18]. The refresh path opens two minutes early, because the fast path only hands back the cached access token while `expires_at` minus a 120-second skew is still in the future [3]. That gives a 120-second band in which any of those five will independently decide the token is stale [15]. The overlap in the worked example is 40 ms wide [2], roughly three hundredths of a percent of the band [16]. A single-process dev box never rolls that dice.
Retrying is the wrong layer, and the writeup is blunt about the reason: a retry loop turns one replayed token into five, which is indistinguishable from the attack reuse detection was built to catch [11]. The damage is also not confined to the loser. Where the provider revokes the family, worker A's brand-new token dies alongside B's replayed one [2], and the recovery path runs through a human pressing Reconnect [20].
The lock itself is the cheap part. An advisory transaction lock keyed on a stable hash of the integration id needs no new infrastructure, has no lease expiry to reason about, and releases when the transaction ends, including when the connection dies [13][4]. The step that does the actual single-flighting is the re-read inside the lock [4]; serialising the workers without re-checking state would queue five refreshes rather than running them concurrently, which is the same number of redemptions arriving in a politer order. The configured 15-second `lock_timeout` is the ceiling on how long a queued worker waits before the statement gives up [4][17].
One caution about the evidence. This is a single writeup with no incident data attached, and it says many providers treat reuse as replay without naming them [20]. It also says the behaviour can differ per app registration, which makes provider documentation the weaker authority [8]. The check is one refresh call: compare the `refresh_token` in the response against the one you sent, and if it came back changed, every refresh from then on is a state mutation you have to serialise [9]. If it came back identical, you are chasing one of the three steady causes instead [6].
What to watch
- Whether providers begin documenting what reuse detection actually revokes, the single token or the whole family, instead of leaving integrators to infer it from a dead account.
- Whether OAuth 2.1 shipping rotation as the default flips providers you currently treat as non-rotating, widening the set of integrations that need serialised refresh.
- Whether advisory transaction locks hold their auto-release guarantee behind a transaction-mode connection pooler, where the transaction boundary is not the application's.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+22
- Incentives
- Insufficient
- Confidence45
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The prescribed fix is to make refresh a single-flight operation per integration and to stop treating access token expiry as something every worker discovers independently.
- [2]
In the worked race, worker A reads the row, sees expires_at in the past and POSTs; worker B reads the same row 40 ms later, sees the same stale expires_at and POSTs the same refresh token; A gets a new pair, B presents an already-redeemed token and gets invalid_grant, and depending on the provider the family is revoked and A's brand-new token is dead too.
- [3]
The sample code sets REFRESH_SKEW to 120 seconds and takes a fast path that returns the stored access token with no lock and no network call when expires_at minus the skew is still in the future.
- [4]
Inside the transaction the sample sets lock_timeout to 15s, calls pg_advisory_xact_lock with a stable 64-bit signed key derived from a SHA-256 hash of the integration id, then re-reads the token row inside the lock because someone may have refreshed while it waited.
- [5]
The OAuth 2.0 spec assigns invalid_grant to any grant that is invalid, expired, revoked, does not match the redirect URI, or was issued to another client, so providers pile several unrelated conditions into one error string.
- [6]
On a refresh call, invalid_grant means one of four things: the refresh token expired through idle-expiry, the user or an admin revoked access, the client credentials do not match the token, or the refresh token was already used once and the provider rotates.
- [7]
Only the already-used-with-rotation case comes and goes; if the error rate is bursty, correlates with traffic spikes or cron minute boundaries, and affects healthy accounts that just worked, the cause is rotation plus concurrency.
- [8]
Refresh-token rotation is a recommended practice in the OAuth 2.0 Security Best Current Practice and the default in OAuth 2.1 drafts, so any provider may rotate; the writeup advises checking the response body rather than the docs, because behaviour sometimes differs per app registration.
- [9]
The tell is that the token response includes a refresh_token field with a value different from the one you sent; that provider rotates, and every refresh is then a state mutation you have to serialise.
- [10]
Locally you run one process; in production you run a web dyno, three queue workers and a scheduled job, all sharing one row in the integrations table.
- [11]
A retry loop makes it worse: it turns one replayed token into five, which looks exactly like the attack that reuse detection exists to catch.
- [12]
The window is not when the token expires, it is when several workers first notice.
- [13]
If you already run Postgres, advisory locks are described as the cheapest correct answer: no extra infrastructure, no lease expiry to reason about, and the lock is released automatically when the transaction ends, including on a crashed connection.
- [14]
The suggested oauth_tokens table is keyed by integration_id and stores access_token, refresh_token, prev_refresh_token, expires_at and rotated_at.
- [15]
With a 120-second refresh skew, there is a 120-second band before expiry in which every worker that touches the integration leaves the fast path and heads for the token endpoint.
- [16]
The 40 ms gap between the two workers in the example is about 0.03 percent of the 120-second contention band.
- [17]
The configured lock_timeout of 15 seconds bounds how long a worker queued behind the advisory lock waits before its locking statement aborts.
- [18]
The production topology described puts five processes in contention for the same integration row.
- [19]
When a third-party OAuth integration works for days then fails with {"error":"invalid_grant"}, the cause is usually not clock skew, a wrong client secret or expired consent, but two of your own processes calling the token endpoint with the same refresh token at the same time.
- [20]
When the provider rotates refresh tokens, the second call presents a token that was already consumed, and many providers treat that as replay and revoke the whole token family, which is why the user has to click Reconnect instead of just retrying.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toWhy Your OAuth Integration Randomly Returns invalid_grant (and How to Stop Two Workers From Racing)
1 article · August 24, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.