Product1 publisher3 min readPublished
Three tools, three MRR numbers: the fix is a metric definition, not a better model
PostHog asked Claude, Cursor and its own assistant the same revenue question and got three answers. Its remedy was a governed catalog of definitions stored as ordinary SQL tables.
The Product Desk · Product desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Asking Claude, Cursor, and PostHog AI the same question, "what was our MRR last month?", produced three different queries and three different numbers.
- One tool summed a Stripe table, one found a slightly different Stripe table, and one tried to reconstruct recurring revenue from raw events and got the proration wrong.
- Every method and number was plausible, but there was no way to tell which was right.
- What "MRR" means at PostHog, which table holds it, and how it is calculated all lived in people's heads, so every agent session reinvented the definition from scratch, slightly differently.
- Humans have the same problem: every new analyst needs to learn which revenue table is the real one and how Stripe is connected.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
PostHog put the same question to three AI tools, Claude, Cursor and its own PostHog AI, asking what MRR was last month, and got three different queries and three different numbers [1]. The interesting part is the diagnosis: according to the company's own account, the models were not the problem, the absence of a written definition was [4][8].
The failure modes are specific. One tool summed a Stripe table, a second found a slightly different Stripe table, and the third tried to reconstruct recurring revenue from raw events and got the proration wrong [2]. Every method and number looked plausible, and PostHog says there was no way to tell which was right [3]. Zero of the three agreed [22].
Three things at PostHog existed only as tribal memory: what a metric actually is, which tables to trust and which to avoid, and how sources join [11]. The company notes that a mature project imports dozens of sources and builds hundreds of data models, plenty of which could plausibly answer "revenue", while only one is current and blessed by finance [12]. The join logic is worse, because it is undocumented procedure: a Stripe customer ID maps to an organization property only after it is reformatted, and nothing records that except the analyst who worked it out last time [13]. Human analysts hit the same wall [6]. The difference, PostHog argues, is that agents answer confidently and fill gaps by hallucinating, and nobody thinks to check [7].
The remedy is described as a dictionary of definitions that both humans and agents read from: define a metric once, approve it once, and subsequent queries return the same number [9]. It does not copy data, replace the warehouse or move rows; it describes what is already there [10]. PostHog's own definition is a governed catalog that tells people and machines what each metric is, which tables to trust, and how sources connect [15], where "governed" means nothing becomes official until a human approves it [16].
The implementation detail worth copying is that the catalog is just SQL [17]. Definitions appear as ordinary tables, metrics are a table, and there is no bespoke catalog API for an agent to learn; anything that can execute SQL can already read the whole layer, so discovery is a query rather than an integration [17][18]. The intended behaviour change is small and mechanical: the agent's first move is to check whether an approved metric exists, and if it does, it runs the governed definition instead of writing its own [19]. In PostHog's architecture this sits inside a context warehouse that already pools product events, imported sources such as Stripe, and data models [20].
Two caveats. This is a vendor writing about the product it is building, with no published accuracy measurements, and PostHog's small-company hedge is honest: if you interact with a handful of tables you probably know where everything is [14].
Watch the approval queue. Governance means a human gates every definition [16], and PostHog also says agents will happily draft metric definitions when pointed at a schema [21], which puts drafting capacity well ahead of review capacity. Also watch whether agents check the catalog first in practice rather than reaching for their own SQL [19], and whether undocumented join logic like the reformatted Stripe key ever makes it into the catalog at all [13].