Build1 publisher3 min readPublished
Say "reliably" in a webhook spec and you have bought the whole distributed systems curriculum
One adverb in a requirement turned a POST forwarder into retries, a dead-letter queue, per-endpoint ordering and two layers of semaphore. The author's own tests kept correcting him.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- The author built webhook-delivery, a service that accepts webhook events over HTTP and reliably delivers them to customer endpoints.
- The project was the author's first with both Go and Kafka, started after he got tired of reading about them.
- The stated requirements included retries, exponential backoff, a dead-letter queue, per-endpoint ordering, idempotency, concurrency limits and per-host circuit breaking.
- The requirement list names seven distinct mechanisms.
- The author writes that on the surface the system is "accept a POST, forward a POST", but saying the word "reliably" means you inherit a whole distributed systems curriculum for free.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A developer published a writeup of webhook-delivery, a Go and Kafka service that accepts webhook events over HTTP and delivers them to customer endpoints, built as his first project with either technology [1][2]. The useful part is not the code but the accounting: one adverb in the requirement generated retries, exponential backoff, a dead-letter queue, per-endpoint ordering, idempotency, concurrency limits and per-host circuit breaking [3], which is seven mechanisms hanging off one word [4].
On the surface, he writes, it is "accept a POST, forward a POST", but say "reliably" and you inherit a distributed systems curriculum for free [5]. The five questions he lists are the ones any delivery system has to answer: what happens when the customer endpoint is down, what happens when it is slow rather than down, what happens when two events for the same customer arrive out of order, what happens when your own process crashes mid-delivery, and how you distinguish a failed delivery from a response that got lost [6]. He says none of them have a clean answer [7].
The resulting shape is conventional and worth restating because most teams get it wrong in the details: the API validates, publishes to a Kafka topic called events and returns 202; delivery workers consume that topic, group messages by an orderingKey so one customer's events stay in order, and POST them out; failures go to a retries topic with exponential backoff; permanent failures and events past a maximum age go straight to the dead-letter queue [8]. That paragraph, he notes, took about three weeks to actually get right [9].
The correction that matters is small. A `chan struct{}` with capacity N is a free concurrency limiter, and he needed two of them, a global cap on in-flight deliveries and a per-host cap so one flaky endpoint could not eat the whole pool, both the same primitive [10]. What did not click until he wrote a test is that if the acquire helper drops its `ctx.Done()` branch, a goroutine can block forever on a saturated semaphore even after the parent context is cancelled, because nothing wakes it [11]. The test that pins this, TestDeliverGroupUnblocksFromSaturatedSemaphore, fills the host semaphore, cancels the context, and fails if delivery is still stuck after two seconds [12]. His summary: a blocking channel operation without a `ctx.Done()` escape hatch is not a semaphore, it is a deadlock waiting for a bad day [13].
The other two lessons are about discipline rather than discovery. The delivery path nests a per-host semaphore, a global semaphore and a mutex over circuit breaker state, and writing that as flat Lock/Unlock with early returns is how you leak a lock, so each critical section became a function literal to give `defer` a clean scope [14]. And the worker depends on a one-method Publisher interface rather than a concrete Kafka producer, so tests use a recording publisher that appends to a slice, which is why 45 tests run in about three seconds with the race detector on and no Docker container involved [15], roughly 67 milliseconds per test [16]. Failure classification uses `errors.As` rather than string matching so it can see through wrapped errors [17].
What to watch: the post flags its Kafka section as the meatier part [18] and says it was the system, or the benchmark numbers, that corrected his mental model [19]. The material in hand stops at the error classification helper, before those numbers appear [20], so the two-layer semaphore and the cancellation escape hatch are what is currently transferable.