Skip to content

Build1 publisher3 min readPublished Updated

Metrics live in RAM: why one observability pipeline hides three different failure modes

A dev.to writeup splits observability into three emitters with almost nothing in common. The split explains why cost estimates, flush windows and missing traces keep surprising teams.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Metrics live in RAM: why one observability pipeline hides three different failure modes
Generated illustration

What happened

  • A dev.to post titled "Observability - A Counter in RAM, an ID in a Header, and a Batch Export" says the author's prior model (import an SDK, sprinkle calls, each call fires data to a server, dashboard reads it back; logging with extra steps) is wrong, and that the difference from logging is not in the analysis but in the emission.
  • The post concludes observability is not one system but three mechanically different emitters that share a pipeline: a number in RAM sampled on a timer; an event enriched with an ID and shipped; an ID in a header propagated hop by hop and reassembled later.
  • A metric is not a record you write; calling increment or record changes a number in the app's memory, and nothing is sent when the line runs.
  • Periodically, every 15 seconds for example, either a backend scrapes an endpoint the app exposes or a collector ships the current values out.
  • Metrics are cheap because a million requests is one counter reading "1,000,000", not a million records.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A writeup published on dev.to argues that the common mental model of observability, an SDK you sprinkle through the code so a dashboard can read it back, is wrong in a specific way: the difference from logging is not in the analysis, it is in the emission [1]. The author's conclusion is that observability is not one system but three mechanically different emitters that happen to share a pipeline: a number in RAM sampled on a timer, an event enriched with an ID and shipped, and an ID in a header propagated hop by hop and reassembled later [2]. That matters operationally because each of the three has its own cost curve and its own way of losing data, and a team that models all three as logging with analysis bolted on will price and debug them wrong.

Start with metrics. When you call increment on a counter, nothing is sent; the number changes in your process memory [3]. Periodically, say every 15 seconds, either a backend scrapes an endpoint your app exposes or a collector ships the current values [4]. That is the whole cost argument: a million requests is one counter reading 1,000,000, not a million records [5], a ratio of a million to one in emitted items [6]. It is also why latency percentiles are not a log-parsing problem. According to the post you could never reconstruct a clean p99 graph from log text, because the histogram was built for it at write time [7].

The same mechanism sets the data loss window. The in-memory counter is disposable by design: each instance flushes on a schedule, on serverless a sidecar collector performs a final flush at shutdown, and the backend sums across instances [8]. Follow that through and the exposure is bounded but real. On a 15 second interval, up to 15 seconds of increments exist only in RAM at any moment, and an instance that dies without flushing takes them with it [9]. That is a different kind of gap from a dropped log line, and it does not appear as an error anywhere.

One caution on the cost claim. The post's arithmetic is about request volume collapsing into a single series; it does not address metric label cardinality [10]. Read "absurdly cheap" as a statement about the counter, not a budget guarantee for however many label combinations you attach to it.

Logs behave the way most people assume everything behaves: an event, written out, shipped, searched. The only upgrade is one automatically injected field, the trace ID [11]. Traces are thinner still. When service A calls service B, the SDK puts the trace ID into the outgoing request: an HTTP header called traceparent, a W3C standard, gRPC metadata, or message headers on a queue [12]. Service B records that it belongs to that trace with a parent span, each service exports its spans independently, and the backend matches parent and child IDs to rebuild the tree [13]. There is no agent watching the network and no distributed coordination [14].

That is where the irreversibility sits. If the trace ID was never propagated at request time, no later analysis can invent it; the structure has to exist when the data is born [15]. The author's framing of OpenTelemetry follows from this: not a tool, a treaty, an agreement across languages, frameworks and vendors on what to emit and what to call it [16].

Worth checking this week: whether your shutdown path actually flushes, and whether trace context survives your queue hop.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories