Build1 publisher3 min readPublished
Scale Kafka sinks on lag, not CPU: KEDA to zero, then fix the per-pod drain rate
A lab writeup on dev.to turns always-on Kafka sinks into on-demand workers. The interesting part is not the zero, it is the ceiling that the partition count puts on your spike response.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The author runs a fleet of Kafka sinks: consumer services that read change events from Kafka, apply business logic, and write the result into a service-local database as a query-friendly materialized view.
- Traffic is not steady: changes arrive in bursts, usually from nightly imports or CDC jobs, and the topic is quiet the rest of the day; the sinks are idle roughly 22 hours a day.
- Idle waste: when the topic is quiet each sink still polls Kafka, holds connections, emits metrics, and occupies CPU and memory; multiplied across dozens of sinks and several regions you pay around the clock for work that happens for a couple of hours a night.
- Spike lag: when the burst lands a backlog builds fast, and if consumers cannot drain it quickly enough consumer lag climbs and downstream reads start serving stale data.
- Sink work is I/O-bound - the consumer waits on Kafka polls and database writes rather than burning CPU - so during a backlog CPU stays flat while lag climbs, and an HPA watching CPU concludes everything is fine and never scales.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
An engineer writing on dev.to has published a measured account of moving a fleet of Kafka sinks from always-on replicas to KEDA-managed scale-to-zero, driven by consumer-group lag rather than CPU [7][8]. The author states every number in the post comes from a local lab that readers can rerun, with the code on GitHub and the commands in an appendix [13]. That matters because the failure mode it documents is the default configuration in most clusters: an HPA watching CPU on an I/O-bound consumer.
The workload shape is the whole argument. These sinks read change events, apply business logic, and write a query-friendly materialized view into a service-local database [1]. Traffic arrives in bursts from nightly imports or CDC jobs, and the topic is quiet the rest of the day, with sinks idle roughly 22 hours out of 24 [2]. That is about 92 percent of the day spent polling Kafka, holding connections, emitting metrics, and occupying CPU and memory for nothing [15][3]. One small sink is cheap; the author's point is that dozens of them across several regions are not [3].
The reflex fix makes the spike worse, not better. Sink work is I/O-bound, so when a backlog builds, CPU stays flat while lag climbs, and an HPA watching CPU concludes everything is fine and never scales [5]. Lag grows silently and downstream reads start serving stale data [4]. Consumer lag is the signal that actually answers the question the autoscaler is being asked, which is whether work is waiting [6].
KEDA supplies that. A plain HPA cannot go from one replica to zero; KEDA manages the HPA and the scale-to-zero around it, and can take Kafka consumer-group lag as the trigger [7]. The published ScaledObject sets minReplicaCount 0, maxReplicaCount 12 to match the partition count, cooldownPeriod 60, lagThreshold 1000, and activationLagThreshold 0 [8]. The last two are the tuning that decides behaviour: activation at zero wakes the workload the instant lag appears, and the cooldown keeps it up until the backlog is fully drained [17]. Replicas then track roughly totalLag divided by lagThreshold, capped at the maximum [9]. Idle, the sink sits at zero pods; a one-million-message burst produced a 0 to 12 to 0 cycle [10].
Now the arithmetic that the config quietly hides. One million messages against a lagThreshold of 1000 asks for 1,000 replicas, and the cap grants 12 [16]. The lag math is not sizing your response during a real burst; the partition count is [8][16]. Above that ceiling the only remaining lever is how fast a single pod drains, which is set by consumer configuration, not by KEDA - a fleet of slow pods still drains slowly [12].
The author is straight about the limits of the run. All 12 pods wrote to one shared database in the lab, which capped aggregate drain rate, so the result is not 12 times one pod; per-pod throughput is a separate measurement [11]. Scaling the write path is deferred to a follow-up [18], and the post's third move, keeping autoscaling from sabotaging itself, is presented as the remaining piece [14].
What to watch: whether the per-pod throughput numbers land with the same lab reproducibility, and whether the write path holds when the shared database is removed as the bottleneck [11][18]. Until then, treat 12 replicas as a partition-bound ceiling and measure one pod before you trust the fleet.