Build1 publisher3 min readPublished
Microsoft's Limbo tests catch AI agents duplicating writes in up to 74% of in-flight and redelivery faults
Microsoft researcher Jiapeng Li's Limbo benchmark found AI agents duplicated side-effecting writes in 56% and 74% of in-flight and double-delivery episodes. Offering idempotency keys cut duplicates from 28% to 4%, so tools that move money or send mail should require them.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- In 90% of the episodes where an agent had duplicated an effect, the agent reported success.
- When a quick read-back could reveal the outcome, frontier models duplicated a write whose acknowledgement was lost in only 0.5% of episodes.
- Limbo graded 25,930 episodes across nine models and three production agent harnesses, checking six simulated services with injected faults against a ledger of what committed.
- The IETF Idempotency-Key header draft reached revision 07 in October 2025, and that revision expired on April 18, 2026, without becoming a standard.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- exposure An agent's own success message cannot be the audit record for writes. Finding double charges or duplicate emails means reconciling against the service's ledger, the way Limbo graded its episodes.
- decision Harness authors have to choose where keys are created. A key generated inside the retry loop changes on every attempt, and the server then has no way to match a retry to the original call.
- constraint Without a finished standard header, each tool integration defines its own key scope, retention and mismatch errors, so a harness cannot assume one convention across services.
- cost Read-back alone covers only the lost-acknowledgement case. The work of adding key support lands on whoever owns each write service an agent can reach.
The figures come from a dev.to tutorial that summarises the paper and rebuilds its idea as a small TypeScript retry layer against a fake billing service [15]. It models three faults [10]. A lost acknowledgement means the charge commits and the client sees a timeout. With a late commit, the request is still in flight when the client gives up; in the demo it lands 90 seconds later. Redelivery sends one request twice while the client sees a single success [10]. To the client, the first two look identical [11].
That split explains the spread in Li's results [1]. After a lost acknowledgement the charge is already in the ledger, so a model that reads back finds it [3]. After a late commit the same read comes back empty, because the write has not landed yet. A model that trusts the empty read retries, and the original arrives later. Redelivery leaves the model nothing to check. Across those last two cases, the duplicate rates are at least 112 times the read-back rate [2].
Only the server sees both copies of a redelivered request. It can drop the second one only if the request carries a key it has already stored. With a key offered on every write, Limbo's duplicate rate fell to about a seventh of the keyless rate [1]. The write-up does not say what produced the duplicates that remained.
The rules attached to a key matter as much as the key. Stripe saves the status code and body of the first request for a key "regardless of whether it succeeds or fails," and returns an error when a reused key arrives with different parameters [7]. That second rule catches a harness that attaches one key to two different charges. Malcolm Featonby's 2020 Amazon Builders' Library article describes Amazon's preferred version, a unique, caller-provided client request identifier in the API contract [8]. Both were written for retry loops that people wrote, and the models appear to have read the same pages: offered a key, they use it [5]. The tutorial pins one key per intent [12] and describes its design as "the agent proposes, the harness owns what actually happens" [13].
Treat the rates as conditional. Limbo's faults are injected [2], so a 74% duplicate rate is measured given a fault. It becomes a per-call rate in production only if your transport fails as often as the sandbox makes it fail. The behaviour after a fault is the part that transfers. The TypeScript repo does not test that either; its author calls it "my small model of the paper's idea, not the Limbo benchmark," with scripted faults and a fake billing service [14].
In my view the key belongs in the schema of every side-effecting tool an agent can call, with the read-back step kept for the lost-acknowledgement case it already handles [3].
What to watch
- Whether the IETF Idempotency-Key draft is revived past revision 07 and moves toward a published standard.
- A breakdown in the Limbo paper of the residual 4%: models skipping offered keys, keys minted per attempt, or services mishandling them.
- Whether the three production agent harnesses Limbo tested start minting per-intent keys by default.