Skip to content

Build1 publisherNot yet confirmed elsewhere3 min readPublished

S3 to Lambda is async and at-least-once: the 3% that vanished after eight quiet months

A pipeline that ran clean for eight months lost roughly 3% of objects once file sizes grew. The retry budget was three attempts, and nothing had been configured to catch the fourth.

The Engineer · Build desk

How we use AISend a correction

What happened

  • An ingestion pipeline ran clean for eight months, then a partner sent larger batch files and about 3% of objects disappeared with no error, alert or CloudWatch trace.
  • S3 event delivery to Lambda, via bucket notifications or EventBridge, is asynchronous and at-least-once, so duplicate invocations for the same object are normal.
  • Lambda's async layer retries, and once those attempts are spent the event goes to a failure destination or DLQ, or is discarded silently if neither exists.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Three delivery attempts is the entire default forensic record. Without a DLQ there is no artifact to reconstruct which objects were lost, so the post-incident question cannot be answered at all.
  • exposure Missing dedup puts the damage on the money paths first, and finance finds it before engineering does, which means the discovery arrives as a billing dispute rather than an alarm.
  • cost A processed-output prefix pointed at its own trigger converts a one-line mistake into consumed account concurrency and an overnight bill, with every other function in the account paying the...
  • decision Raising on the first bad record hands control of your failure blast radius to S3's batching behaviour; collecting per-record failures takes it back.

The pipeline did not acquire a bug in month eight. It kept doing what it had always done, which was to treat Lambda's automatic retries as error handling, because until then the retries had quietly worked [14]. That holds while failures are transient. Larger files make failures deterministic: whatever the object size crossed, it gets crossed again on attempt two and again on attempt three, and then the event is discarded with no destination configured to receive it [1][2]. The account published on dev.to, originally on kuryzhev.cloud, does not name the limit that broke, and the arithmetic matters more than the limit anyway. One delivery plus two default retries is three chances, all spent on the same input that fails the same way [11]. The configurable range is 0 to 2, so the floor is a single attempt before the event is gone [1][12].

Note where the bytes come from. The event carries bucket, key, size, eTag and versionId, and the function calls GetObject itself [7]. File growth therefore lands inside the invocation, not inside S3's delivery, which is why a partner changing batch size is a change to your function's runtime behaviour and not to your integration.

The silence is the operationally expensive part. CloudWatch shows two failed invocations and then nothing, which reads exactly like a transient problem that resolved itself [4]. Roughly 3% of objects disappeared with no error, no alert and no trace [13], and the shape of that evidence is indistinguishable from success.

Two more assumptions compound it. S3 can pack multiple records into one invocation payload, and does so particularly under load [15], so a handler that reads event['Records'] as a single item starts losing data at exactly the moment volume rises [15]. Wrap that handler in one try/except that raises on the first bad record and five records become five losses when one is malformed, with the other four either reprocessed for nothing or gone if the retries also fail [9]. The blast radius of a single bad object is set by S3's batching decision, not by yours.

The fix in the post is structural rather than clever: iterate records independently and collect failures instead of raising on the first error, described as the single biggest change [5]. The sample pairs that with a conditional write to a DynamoDB table keyed on bucket, key and eTag, so a repeat delivery is a no-op [10]. That guard is not optional hygiene. At-least-once delivery means duplicate invocations for the same object are expected behaviour [6], and without dedup the visible symptom is duplicate rows, double-charged invoices or a revenue report that double-counts an afternoon [3].

Two other settings belong in the same design pass. Reserved concurrency, because a bulk sync of 10,000 files spikes concurrent invocations at once and an unbounded function either throttles the account or drops events it cannot absorb [16]. And the output prefix, because writing processed results back into the prefix that triggers the function loops, exhausts account concurrency in minutes, and produces a bill overnight [8].

The useful question about an S3 trigger is not whether the handler works. It is what artifact exists after it fails.

What to watch

  • The published code sample cuts off inside the idempotency guard, so how the collected failures are actually reported or re-driven is still unverified.
  • Which limit the larger files broke, since timeout, memory and throttling point at different fixes: reserved concurrency, a DLQ, or both.
  • Whether teams move these events onto EventBridge or a queue in front of Lambda to get retention and redrive instead of a fixed three attempts.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence38
Adoption
Insufficient
Hype gap+18
Incentives55
Confidence44
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    S3 does not retry; Lambda's async invocation layer does, with two automatic retries by default, configurable via MaximumRetryAttempts in the range 0 to 2.

  2. [2]

    After the retries, the event goes to a configured failure destination or DLQ; if none is configured it is discarded silently.

  3. [3]

    Without a dedup mechanism, at-least-once delivery produces duplicate database rows, double-charged invoices or duplicate downstream messages, and teams typically notice when someone in finance asks why a report double-counted revenue for one afternoon.

Sources

1 independent publisher whose own reporting we read for this story.

  1. dev.to

    1 article · August 22, 2026

    Fix Lambda S3 Trigger Error Handling Before It Loses Events

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

  • Idempotency PatternsFollow
  • Pipeline reliability and silent data lossFollow
  • Dead-letter queues and failure destinationsFollow
  • Serverless event delivery semanticsFollow
  • Lambda concurrency and cost controlFollow
Loading related stories