Build1 publisherNot yet confirmed elsewhere3 min readPublished
S3 to Lambda is async and at-least-once: the 3% that vanished after eight quiet months
A pipeline that ran clean for eight months lost roughly 3% of objects once file sizes grew. The retry budget was three attempts, and nothing had been configured to catch the fourth.
The Engineer · Build desk
What happened
- An ingestion pipeline ran clean for eight months, then a partner sent larger batch files and about 3% of objects disappeared with no error, alert or CloudWatch trace.
- S3 event delivery to Lambda, via bucket notifications or EventBridge, is asynchronous and at-least-once, so duplicate invocations for the same object are normal.
- Lambda's async layer retries, and once those attempts are spent the event goes to a failure destination or DLQ, or is discarded silently if neither exists.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint Three delivery attempts is the entire default forensic record. Without a DLQ there is no artifact to reconstruct which objects were lost, so the post-incident question cannot be answered at all.
- exposure Missing dedup puts the damage on the money paths first, and finance finds it before engineering does, which means the discovery arrives as a billing dispute rather than an alarm.
- cost A processed-output prefix pointed at its own trigger converts a one-line mistake into consumed account concurrency and an overnight bill, with every other function in the account paying the...
- decision Raising on the first bad record hands control of your failure blast radius to S3's batching behaviour; collecting per-record failures takes it back.
The pipeline did not acquire a bug in month eight. It kept doing what it had always done, which was to treat Lambda's automatic retries as error handling, because until then the retries had quietly worked [14]. That holds while failures are transient. Larger files make failures deterministic: whatever the object size crossed, it gets crossed again on attempt two and again on attempt three, and then the event is discarded with no destination configured to receive it [1][2]. The account published on dev.to, originally on kuryzhev.cloud, does not name the limit that broke, and the arithmetic matters more than the limit anyway. One delivery plus two default retries is three chances, all spent on the same input that fails the same way [11]. The configurable range is 0 to 2, so the floor is a single attempt before the event is gone [1][12].
Note where the bytes come from. The event carries bucket, key, size, eTag and versionId, and the function calls GetObject itself [7]. File growth therefore lands inside the invocation, not inside S3's delivery, which is why a partner changing batch size is a change to your function's runtime behaviour and not to your integration.
The silence is the operationally expensive part. CloudWatch shows two failed invocations and then nothing, which reads exactly like a transient problem that resolved itself [4]. Roughly 3% of objects disappeared with no error, no alert and no trace [13], and the shape of that evidence is indistinguishable from success.
Two more assumptions compound it. S3 can pack multiple records into one invocation payload, and does so particularly under load [15], so a handler that reads event['Records'] as a single item starts losing data at exactly the moment volume rises [15]. Wrap that handler in one try/except that raises on the first bad record and five records become five losses when one is malformed, with the other four either reprocessed for nothing or gone if the retries also fail [9]. The blast radius of a single bad object is set by S3's batching decision, not by yours.
The fix in the post is structural rather than clever: iterate records independently and collect failures instead of raising on the first error, described as the single biggest change [5]. The sample pairs that with a conditional write to a DynamoDB table keyed on bucket, key and eTag, so a repeat delivery is a no-op [10]. That guard is not optional hygiene. At-least-once delivery means duplicate invocations for the same object are expected behaviour [6], and without dedup the visible symptom is duplicate rows, double-charged invoices or a revenue report that double-counts an afternoon [3].
Two other settings belong in the same design pass. Reserved concurrency, because a bulk sync of 10,000 files spikes concurrent invocations at once and an unbounded function either throttles the account or drops events it cannot absorb [16]. And the output prefix, because writing processed results back into the prefix that triggers the function loops, exhausts account concurrency in minutes, and produces a bill overnight [8].
The useful question about an S3 trigger is not whether the handler works. It is what artifact exists after it fails.
What to watch
- The published code sample cuts off inside the idempotency guard, so how the collected failures are actually reported or re-driven is still unverified.
- Which limit the larger files broke, since timeout, memory and throttling point at different fixes: reserved concurrency, a DLQ, or both.
- Whether teams move these events onto EventBridge or a queue in front of Lambda to get retention and redrive instead of a fixed three attempts.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+18
- Incentives55
- Confidence44
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
S3 does not retry; Lambda's async invocation layer does, with two automatic retries by default, configurable via MaximumRetryAttempts in the range 0 to 2.
- [2]
After the retries, the event goes to a configured failure destination or DLQ; if none is configured it is discarded silently.
- [3]
Without a dedup mechanism, at-least-once delivery produces duplicate database rows, double-charged invoices or duplicate downstream messages, and teams typically notice when someone in finance asks why a report double-counted revenue for one afternoon.
- [4]
With no DLQ, CloudWatch shows two failed invocations and then nothing, and teams assume the problem resolved itself.
- [5]
The recommended fix is to structure the handler to iterate records independently and collect failures instead of raising on the first error, described as the single biggest change that fixes most of the listed problems.
- [6]
S3 event notification delivery, whether via the bucket's notification configuration or EventBridge, is asynchronous and at-least-once; S3 does not wait for the Lambda to finish and does not guarantee exactly-once delivery, so duplicate invocations for the same object are expected behaviour.
- [7]
Lambda does not receive the file; it receives metadata including bucket name, object key, size, eTag and versionId if versioning is on, and the function must call GetObject itself to get the bytes.
- [8]
Writing processed output back into the same prefix that triggers the function causes recursive invocation, which can burn through account concurrency limits in minutes and generate a surprising AWS bill overnight.
- [9]
Wrapping the whole handler in one try/except and raising on the first bad record means a batch of five with one malformed record loses all five; the good four are reprocessed unnecessarily or lost if retries also fail.
- [10]
The sample handler uses a DynamoDB table named processed-objects with a partition key of bucket/key/eTag and a conditional put_item as an idempotency guard.
- [11]
Under the default configuration an S3 event gets three delivery attempts in total before it is discarded or routed to a failure destination.
- [12]
Because MaximumRetryAttempts can be set to 0, the minimum configurable delivery budget is a single attempt before the event is discarded or sent to a destination.
- [13]
A client's ingestion pipeline "worked fine" for eight months; then a partner began uploading larger batch files, retries kicked in, and roughly 3% of objects vanished with no error in the logs, no alert and nothing in CloudWatch.
ReportedInsufficientSource: dev.to post, originally published on kuryzhev.cloud2 sources— create a free account to open themView cited source - [14]
The root cause was not a bug in the client's code but S3 trigger error handling nobody had designed, because the trigger had always quietly retried and succeeded before.
- [15]
S3 can batch multiple records into a single invocation payload, especially under load, so a handler that assumes event['Records'] has exactly one item will drop data the first time it does not.
- [16]
Without reserved concurrency, a bulk S3 sync of 10,000 files spikes concurrent invocations instantly, and the function either throttles account-wide or silently drops events Lambda cannot absorb.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toFix Lambda S3 Trigger Error Handling Before It Loses Events
1 article · August 22, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.