Skip to content

Product1 publisher3 min readPublished

One dashboard, two answers: 41,900 units, then 43,100, with no code change

A Tiger Data explainer traces the gap to late-arriving rows, and to the fact that four refresh strategies filed under one label disagree precisely there.

The Product Desk · Product desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying One dashboard, two answers: 41,900 units, then 43,100, with no code change
Generated illustration

What happened

  • A post published on timescale.com opens with a shift dashboard that reported 41,900 units across an eight-hour window one day and 43,100 for the same query over the same eight hours the next, with no code change shipped and no row edited after the fact.
  • The stated cause was a plant-floor historian that had lost its uplink, buffered locally, and flushed overnight, landing rows that carry their original timestamps.
  • The difference between the two reported figures is 1,200 units, about 2.9 percent of 41,900.
  • The raw table was correct at every point along the way; the pre-computed aggregate behind the dashboard never went back for those rows, so it published one number and later replaced it with another.
  • Every pre-computation strategy handles a fresh row arriving at the head of the table.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

A post published on timescale.com opens with a shift dashboard that reported 41,900 units across an eight-hour window one day and 43,100 for the same query over the same eight hours the next, with nobody shipping a code change and nobody editing a row after the fact [1]. The cause, in the post's telling, was a plant-floor historian that lost its uplink, buffered locally and flushed overnight, landing rows that carried their original timestamps [2]. That is a 1,200-unit move, about 2.9 percent of the first figure [3]. The raw table was correct at every point along the way; the pre-computed aggregate behind the dashboard never went back for those rows, so it published one number and later replaced it with another [4]. The useful part of the argument is the framing. Every pre-computation strategy handles a fresh row arriving at the head of the table [5]; they diverge on the two events demos skip, which are data arriving late for a window already computed and data that changes after the fact [6]. The industry files all of them under "materialized view refresh," which implies one correctness story where there are four [7]. Each design bets on when the work happens: at write time, at schedule time, or at query time [8]. Rebuild-on-schedule is the first. In vanilla Postgres, REFRESH MATERIALIZED VIEW rebuilds from scratch with no incremental path, so the view is stale between runs by construction [9], and you pay the full rebuild whether one row changed or a million [10]. That works for as long as a rebuild of the range you care about fits inside the schedule interval [11]. A daily rollup over a month of data on modest hardware fits for a long time; a fine-grained rollup over years of raw samples stops fitting well before anyone notices, because the failure shows up as schedule slip rather than an error [12]. The post puts InfluxDB v2 tasks in the same shape, and notes that the documented lever for lateness there is an offset that delays the run rather than an invalidation layer that reaches back [13]. Insert-triggered incremental views are the second. ClickHouse standard materialized views fire on INSERT into the source table and write the increment forward [14], which is what makes them viable at ingest rates where a recompute is out of the question [15]. The premise is that the past never changes: nothing in the standard view reacts to a backfill or a correction, so reconciling history becomes a manual, resource-heavy operation [16]. ClickHouse's own refreshable materialized views re-run the full query on a schedule and atomically swap the destination table, which is strategy one again, reached for because the premise broke [17]. Third, RisingWave and Materialize maintain a view through a dataflow graph updated on every source change, targeting sub-second freshness, so late data is just another change event and there is no refresh window to reason about [18]. The post concedes the cost is architectural rather than computational: a persistent streaming system alongside your database, with its own consistency model and operational surface [19]. Fourth is the vendor's own. A TimescaleDB continuous aggregate is a query with a time bucket and a GROUP BY whose results live in a table TimescaleDB maintains, kept current by a scheduled policy [20]. Writes that touch already-summarized ranges land in an invalidation log, and each run re-examines a bounded window, recomputing only the buckets inside it that saw activity [21], which keeps the work small and inside the database you already run [22]. What decides the outcome is narrower than the taxonomy. On a continuous aggregate, two offsets determine whether a late row is ever reconciled at all [23].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories