Build1 publisher3 min readPublished
A Prometheus that had written nothing for hours passed every health check
Compaction failed sixty times an hour on a full volume while dashboards stayed correct, because queries are served from memory. Liveness checks cannot see durability.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- agentic-os is a small self-hosted observability stack run on Railway comprising Prometheus, Grafana, an OpenTelemetry Collector, a cloudflared tunnel, and a public status API that publishes three numbers about the author's Claude Code usage: requests, tokens, cost.
- The author describes agentic-os as a personal project but a real deployment with real uptime whose entire purpose is to be the thing that tells him when something is wrong.
- On 13 August, Prometheus stopped writing blocks to disk; its TSDB compaction started failing with "no space left on device" once a minute and never stopped.
- The author characterises sixty compaction failures per hour as not a degradation but every single attempt failing.
- During the failure the public status page kept serving correct, moving numbers, every HTTP health check was green, and Grafana kept drawing graphs.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
On 13 August, a Prometheus instance stopped writing blocks to disk, and its TSDB compaction started failing once a minute with "no space left on device" and never stopped [2]. Sixty failures an hour is not degradation, it is every single attempt [3], and none of it registered on any instrument pointed at the process: the public status page kept serving correct, moving numbers, every HTTP health check was green, and Grafana kept drawing graphs [4].
The stack is agentic-os, a self-hosted deployment on Railway running Prometheus, Grafana, an OpenTelemetry Collector, a cloudflared tunnel, and a public status API that publishes request, token, and cost numbers for the author's Claude Code usage [1]. It is a personal project, but a real deployment whose entire purpose is to say when something is wrong [c1b].
The structural point is the one worth carrying out of this repo. Prometheus answers queries out of the head block, which lives in memory, so from the outside a database that had persisted nothing for hours was indistinguishable from a healthy one [5]. A health check on a TSDB exercises the read path, and the read path was fine. Durability is a separate subsystem that emits its own signals, and in this stack nothing was scraping prometheus_tsdb_* metrics at all, so there was no series to write an alert rule against [9]. The author found the fault by reading container logs by hand [8].
The usual reassurance does not close the gap either. A restart is not instant data loss, because the head is reconstructed from the write-ahead log [6], but the WAL sat on the same volume that had run out of space [7], which makes "it recovers on restart" a bet on the one resource already exhausted.
The remediation was to scrape Prometheus with Prometheus, cap retention by size rather than time alone, and add a watchdog reporting to Sentry [10], driven by the query sum(increase(prometheus_tsdb_compactions_failed_total[1h])) [11].
Then the alerting path produced a worse failure. Sentry received one event, then nothing for six days [12]. Checked on 19 August, the issue read "last seen five days ago", which is what a cleared transient looks like, and the author nearly closed it [13]. The deduplication state was a module-level set: one notification per process lifetime [14]. The four events that did arrive were four restarts, not four detections [15]. Behind that shape, compaction was still failing sixty times an hour, on the order of 8,640 failed attempts across the silent stretch [1].
The author is precise about the defect: deduplication is correct, since a per-minute failure must not produce 1,440 events a day, and the bug was choosing a window of "forever," which cannot distinguish "happened once" from "still happening" [16][17]. The replacement stores a per-key last-sent timestamp with a 3,600 second interval [18], which caps the same condition at roughly 24 events a day instead of 1,440 [2]. The same code path also carried a cost defect, where the error path cost more than the success path [19].
Four things to check on your own stack, in order of how cheap they are: whether anything scrapes the monitoring system's own TSDB metrics, whether retention is bounded in bytes as well as days, whether the WAL shares the volume you are betting a restart on, and what your dedup does after the first event. Readers clicking through to PR #51 and PR #86 should note the repo's commit messages and code comments are in Italian [20].