Product1 publisher3 min readPublished
Four-times-faster incident pipeline caught just 64% of in-scope incidents in a bad month
One SaaS team's rebuild of incident detection on Kafka and Flink cut telemetry lag from over 40 seconds to under 10, according to its CNCF post. Over eighteen months its recall ran from about 60% to 86% and back to 64%, so speed alone says little about which outages get caught.
The Product Desk · Product desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Run cost for the first-generation system climbed from roughly $120K to $230K a year as products and experiences were added.
- Two incidents in one year were caused by another tenant's lag on the shared cloud queue that fed the old Node.js aggregator.
- The rebuild filters an allow-list of products and experiences at the bus onto a dedicated Kafka topic whose seven-day retention doubles as a replay window.
- A single Apache Flink 1.20 job, deployed through the Flink Kubernetes Operator, handles stream processing, and detectors currently run on a commercial metrics service.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- decision Anyone presenting a detection rebuild has to pick a headline figure, and this record puts a one-time latency gain beside a recall series that gave back 22 points.
- constraint Error-rate detection on client telemetry cannot see customers whose pages never load, so a faster pipeline leaves that class of outage uncovered until someone builds volume-drop detection.
- exposure Monitoring that shares transport with other workloads can lose sight of an incident during a neighbour's backlog, so ingest isolation belongs in the detection system's reliability requirements.
A customer on a hard-down database shard tries to open a board, and the page never loads. In this team's setup, client-side telemetry only exists when a page loads, so that customer sends no event at all [7]. The tenant's error-rate chart stays flat. The authors wrote: "A hard-down database shard produces silence, not errors. Any design that only looks for failures will read silence as health." [7]
The headline figure measures one stage of the path. The old 40 seconds was event-to-metric time on a good day [3]. Getting that stage under 10 seconds [2] cuts at least 30 seconds, better than fourfold [1]. After that stage, the old detectors needed several minutes of windowed data before they could fire [3].
Recall is what an on-call lead answers for. Across eighteen months of monthly measurement [1], in-scope recall went from about 60% to a peak of 86%, then fell back to 64% in a bad month [4]. The climb was 26 points and the fall 22, leaving recall four points above where it started [2]. In that month the system did not flag about 36% of the incidents it was meant to catch [3]. Precision, the team wrote, "is still below where we want it." [5] The authors framed the account this way up front: "It is not a success story with a bow on it." [6]
Here's what teams tell themselves users do: hit an error and show up on the error-rate chart. Here's what users on a dead shard actually do: nothing the pipeline can see, at any latency. The post's own definition of something wrong covers both cases: "An error rate or a volume drop that a human would call an incident." [13] Only the volume-drop half reaches the customer on the dead shard.
The post is most useful to teams running central monitoring over client telemetry at volume. This company runs more than ten cloud products for millions of tenants and handles billions of events a day [8]. Its post opens on the question such teams face after a major incident, "who noticed first, the monitoring or the customers?", and the team then says that for a long time its honest answer was that it depended [14].
I think the useful habit is to report recall every month beside latency, and to lead with recall. The cost is labour. Recall needs a list of the incidents the system should have caught, and someone has to rebuild that list every month. The forcing function is a 2x2 over last quarter's high-severity incidents. One axis is who noticed first, monitoring or customer. The other is what the signal looked like, errors or silence. Where customers noticed first and the signal was errors, a faster pipeline and tighter thresholds help. Where customers noticed first and the signal was silence, the fix is a volume-drop detector, and pipeline speed will not move those cases. The monitoring-first half of the grid is what the latency headline describes. I'd expect silence detection to add false pages, and this team already wants its precision higher [5].
What to watch
- Recall for the months after the 64% dip, to show whether the system can hold anything near its 86% peak.
- A published precision figure set against the monthly recall series.
- Whether detectors move off the commercial metrics service onto the Prometheus-compatible store the Flink job already feeds through OpenTelemetry.