Build1 publisherNot yet confirmed elsewhere2 min readPublished
One incident instead of two: what SKIP LOCKED actually buys you
Putting the job queue in the customer database costs no new infrastructure and hands autovacuum a race it loses. The benchmark cited against it caps Postgres near 660 messages a second.
The Engineer · Build desk
What happened
- A dev.to post argues that reserving jobs with SKIP LOCKED buys a queue with no new infrastructure by placing it inside the least replaceable component.
- Each claim, completion and retry is an update, and MVCC writes a new row version every time, so the jobs table churns constantly.
- A May 15, 2023 benchmark put Postgres-as-a-queue at roughly 660 messages a second on a 1KB payload with 38ms P95 publish latency.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint The two systems share a recovery path, so the queue cannot drain while the database is thrashing and each retry wave extends the outage instead of clearing it.
- cost Worker polling is paid for out of the same connection pool the product uses, so the bill arrives as customer-facing latency before it shows up in queue metrics.
- decision The sizing question stops being whether SKIP LOCKED works and becomes whether the jobs table's write rate stays inside what autovacuum can keep up with.
- contradiction The same piece that calls this a single point of failure endorses it for low-rate background work, which turns the argument from an architectural ban into an unquantified threshold.
The unit of accounting here is the dead tuple. Claiming a job is an update, completing it is an update, retrying it is another, and Postgres writes a new row version for each one [2]. That makes the jobs table the highest-churn object in a schema whose other tables are mostly read. At the ceiling the cited benchmark reports, about 660 messages a second [5], even a minimal two updates per job puts roughly 1,320 dead row versions a second in front of autovacuum [14]. Gunnar Morling's argument, in his November 3, 2025 piece, is that vacuum eventually loses that race and the WAL piles up behind it [8]. Brandur Leach, formerly a staff engineer at Stripe, describes the same end state as table bloat, index fragmentation and autovacuum starved of room [7]. The vacuum budget and the WAL are not queue resources. They are the customer's.
The throughput gap is the headline number, and it is the less interesting one. RabbitMQ moved 25,000 messages a second in the same test environment [9], about 38 times as much [12]. The comparison worth sitting with is FIFO SQS, the deliberately capped managed option at 3,000 messages a second [10]: still roughly four and a half times the Postgres figure [13]. The complexity being avoided is a queue URL.
Recovery is where the shared substrate actually bills you. AWS engineers described the loop in a December 17, 2021 Architecture Blog post: a stalled delivery process pushes backpressure onto the database, which produces more failed work [11]. A broker backlog is drained by adding consumers. A bloated jobs table is drained by vacuum, which needs the write rate to fall, which needs the workers to back off, which is the opposite of what a retry loop does. Meanwhile every worker node holds a connection, and Postgres forks an OS process per connection, so polling workers churn through the pool that user-facing queries are drawing from [3].
The evidence for the throughput half of this case is thin: one benchmark, one date, one payload size, 38ms at P95 [5], with no methodology quoted. Hardware and batching would move that figure a long way. The coupling half needs no benchmark, and the article's own carve-out is the useful sentence in it: sporadic background work at low rates is fine, and the failure is assuming that behaviour survives production volume [4]. So the threshold is a metric, not an architecture. Watch dead tuples and vacuum lag on the jobs table against the rest of the schema, and answer the question the piece poses directly, which is whether your database goes down when your queue does [6].
What to watch
- Whether the May 15, 2023 benchmark's setup is published: hardware, batching and index layout decide how much weight the 660 per second figure can carry.
- Whether teams reporting success run workers behind a connection pooler, which blunts the process-per-connection argument while leaving the bloat argument intact.
- Dead tuple counts and vacuum lag on the jobs table against the rest of the schema, which is the metric that says when a working queue stops working.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+34
- Incentives38
- Confidence36
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
SELECT ... FOR UPDATE SKIP LOCKED lets a team reserve jobs with zero new infrastructure: no brokers, no ops, one line of code, hidden inside the most critical component.
- [2]
Postgres MVCC creates a new row version for every update, and a queue is a table that is repeatedly modified: claiming a job updates it, completing it updates it, retrying it updates it again.
- [3]
Postgres creates a new OS process for each connection, and many worker nodes constantly polling, claiming and updating jobs cause continuous connection churn that can deplete the pool, making user-facing queries compete with job runners.
- [4]
The author does not rule the pattern out: Postgres queues are fine for minor internal tasks or a sporadic background job at low job rates, and the failure is assuming that what works at low volume works at scale with production traffic.
- [5]
A benchmark on May 15, 2023 indicated Postgres-as-a-queue maxed out at approximately 660 messages per second with a 1KB payload and a 38ms P95 publish latency.
- [6]
The author's reality check: if your database would go down when your queue does, you did not build a shortcut, you built a single point of failure.
- [7]
Brandur Leach, who used to be a staff engineer at Stripe, detailed the result in his article 'Transactionally Staged Job Drains in Postgres': table bloat, index fragmentation, and autovacuum starved for air.
- [8]
Gunnar Morling argued in his November 3, 2025 analysis "'You Don't Need Kafka, Just Use Postgres' Considered Harmful" that long-running consumer transactions inevitably lead to MVCC bloat and WAL pile-up, and that vacuum loses the race against the change rate.
- [9]
RabbitMQ processed 25,000 messages per second in the same test environment.
- [10]
Standard Amazon SQS provides high throughput almost without limit, and FIFO SQS achieves exactly-once processing with a ceiling of 3,000 messages per second, or 30,000 in a batch.
- [11]
AWS engineers warned in a December 17, 2021 Architecture Blog post that an inoperable delivery process places backpressure on the database, creating a feedback loop that results in even more failed work.
- [12]
RabbitMQ's measured rate is about 38 times the Postgres-as-a-queue rate in the same test environment.
- [13]
FIFO SQS's capped 3,000 messages per second is about 4.5 times the benchmarked Postgres-as-a-queue ceiling.
- [14]
At 660 messages per second, with a minimum of two updates per job (claim and completion), the jobs table generates at least about 1,320 dead row versions per second for autovacuum to reclaim.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toPostgres as your message queue is a SPOF you'll regret at 3am
1 article · August 22, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
Entities
- PostgreSQLFollow
- SELECT ... FOR UPDATE SKIP LOCKEDFollow
- RabbitMQFollow
- Amazon SQSFollow
- Apache KafkaFollow
- Brandur LeachFollow
- Gunnar MorlingFollow
- StripeFollow
- Amazon Web ServicesFollow
- dev.toFollow