Build1 publisher3 min readPublished
A Timed-Out Reset SMS Is Not A Failed One, And Your Retry Code Probably Disagrees
Durable admission with a separate event_id and idempotency_key turns an unknown provider outcome into a reconcilable record, and gives auditors evidence without ever storing the token.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Treat an SMS timeout as an unknown outcome rather than a failed send: accept each password-reset event once, persist its expiry and idempotency key before dispatch, and retry only through a worker that can reconcile the original attempt.
- For a short-lived e-commerce reset token, compliance evidence is the deciding constraint: the system must show what it accepted, what it attempted, when it stopped and why, without storing the token or message body in an audit log.
- A Node.js Express handler may receive the event but should not hold the HTTP request open while an SMS provider decides the final delivery state; return an accepted response after durable admission, then expose status from local state.
- Use two identifiers with different jobs: event_id identifies the business action, such as one password-reset request, and idempotency_key identifies the logical notification command.
- A unique constraint on the idempotency key makes two concurrent HTTP requests converge on one stored record; checking memory before an insert is not enough because two processes can pass that check together.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A password-reset SMS that times out has not failed; it has returned an unknown outcome, and a dev.to write-up by Liam Foster argues the correct response is to accept the event once, persist its expiry and idempotency key before dispatch, and retry only through a worker that can reconcile the original attempt [1]. That framing matters because for a short-lived e-commerce reset token the deciding constraint is compliance evidence: the system has to show what it accepted, what it attempted, when it stopped and why, without storing the token or the message body in an audit log [2].
Start with the arithmetic of a timeout. Three realities are consistent with it: the provider never accepted the request, it accepted and the response was lost, or it accepted and sent the message before the caller gave up waiting [6]. Two of those three involve a message that may already be on its way, so an immediate blind retry is the right guess in only one case out of three [19]. Retrying as if the first case were certain is exactly how a customer gets two reset messages, and declaring success instead is no better [7]. The durable record should enter dispatch_unknown, keep the provider's attempt identifier when one exists, and pass through reconciliation before any further send is authorized [8].
The identifiers do two different jobs. event_id names the business action, one password-reset request; idempotency_key names the logical notification command [4]. A unique constraint on that key is what makes two concurrent HTTP requests converge on a single stored record, because an in-memory check before the insert can be passed by two processes at the same time [5].
This reshapes the endpoint. An Express handler can receive the event, but it should not hold the request open while the provider decides a final delivery state; return an accepted response after durable admission and expose status from local state [3]. Polling is not retry: it reads the provider's view and updates the local record, and it must not create a second message, a separation that is easy to blur once a generic checkAndRetry() owns both operations [9].
Expiry is a dispatch boundary rather than presentation metadata. Compare the clock to expires_at before every attempt, and once the token is too close to expiry for a useful delivery, mark the notification expired and stop [10]. Foster declines to name a universal safety margin, arguing it should be resolved from your own latency distribution and product policy and then recorded as auditable configuration [11].
The evidence is specific: each transition needs a timestamp, old and new state, event ID, attempt number and reason code, with credentials, the reset URL, the token and the full phone number kept out [13]. A redacted destination fingerprint can support correlation, but it sits under the same retention and authorization controls as the rest of the record [14]. The contract, then, is admission exactly once for a stable event identifier, no dispatch after expiry, retained state transitions, and inspection that does not trigger work; "exactly once" describes admission of the logical command, not a fiction in which the SMS network joins your database transaction [12].
Watch the build decision, because it is a controls question. A managed notification service owns provider reconciliation and channel routing and cuts on-call surface, but its status vocabulary and evidence export may not match what an auditor expects [15]. Direct integration exposes more provider detail and fewer translation layers, and hands your team leases, retry classification, retention and every 02:00 alert [16]. Self-hosting gives the strongest control over data placement and change timing, and is unsuitable when the team cannot staff queue, database and delivery integration as an on-call product [17]. No option wins by default [18].