Skip to content

Product1 publisher3 min readPublished

A 2x LLM bill is not a bug report: token spend is an observability problem

PostHog says one product's LLM cost doubled from $5k to $10k in a day and it was not a regression. Without per-workflow tracking, growth and a defect look identical on the invoice.

The Product Desk · Product desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • A product's LLM costs doubling from $5k to $10k in a day was not a regression: a launch caused a workflow to run 4x more, but the cost per run fell from $1.8 to $1.4.
  • The Claude Code SDK in PostHog Desktop uses Haiku extensively (3-5 calls per Opus call), revealing an opportunity to swap it with a cheaper, self-hosted model in future.
  • RTK was not saving tokens on bash commands as expected because it was not being used properly by Claude; a fix cut 13% in bash token usage and a PR is open to ship it to everyone.
  • PostHog says teams should monitor three areas: flat-rate subscriptions such as Claude Code that are constrained by context windows and rate limits rather than costs; team use of automations, Slack apps and scheduled agents; and customer use of AI-powered features, where efficiency affects what you can build, charge and grow.
  • Teams almost always have a gap in at least one of these monitoring areas, which limits detail about which workflows, features and use cases consume the most tokens; the monthly bill from Anthropic will not tell you this.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

PostHog's engineering newsletter published a guide to cutting token spend, and the most useful part is not the advice but three diagnoses: a doubled bill that turned out to be growth, a hidden Haiku multiplier, and a tool that was not being called the way its authors assumed [1][2][3]. Each was legible only because someone was measuring tokens per workflow rather than per invoice [5].

The headline case: one product's LLM cost went from $5k to $10k in a day, and it was not a regression [1]. A launch pushed one workflow to run 4x more often while cost per run fell from $1.80 to $1.40 [1], a roughly 22 percent improvement in unit economics buried inside a number that reads as an incident [15]. The two figures do not reconcile cleanly: 4x volume at 78 percent of the old unit cost implies about a 3.1x bill, not 2x [16], which suggests the measurements cover different windows. The direction of the lesson survives the arithmetic. The aggregate is the least informative view available.

The other two findings are smaller and harder to see. PostHog Desktop's use of the Claude Code SDK makes 3 to 5 Haiku calls per Opus call, which the team reads as an opening for a cheaper self-hosted model [2]. And RTK was not saving tokens on bash commands because Claude was not using it properly; the fix cut bash token usage by 13 percent, with a PR open to ship it more widely [3]. Neither shows up in a monthly total.

The coverage argument is the operational one. There are three places to instrument: flat-rate seats such as Claude Code, which are constrained by context windows and rate limits rather than dollars; internal automations, Slack apps and scheduled agents; and customer use of AI features [4]. Teams almost always have a gap in at least one, and the monthly bill from Anthropic will not close it [5]. The tooling named is unglamorous: the /usage command, ccusage, quota widgets like ccseva, and spend monitoring at the gateway [6]. Cost on its own is a trap, so PostHog pairs it with accuracy, success rates and usage, partly because quality tracking surfaces bad queries, repeated retries and agent runs that produce nothing, which is spend that should never have been incurred [7][8].

The second budget is the context window. MCP tool definitions no longer load into the system prompt, but fetching one can still consume a session [9]. The official Atlassian server runs about 10k tokens for Jira and Confluence tools alone [10]; the official GitHub server exposes 94 tools at about 17.6k [11], roughly 187 tokens per tool [17]. Together that is about 27.6k tokens before any work starts [18]. Cloudflare's native MCP with full schemas reaches 2,594 tools and 1,170,523 tokens, which Code Mode fits into 1,069, about a thousandfold reduction [12][19]. PostHog's own server once exposed 183 tools at 113,843 tokens, and no longer does [13].

On AGENTS.md, PostHog cites one study finding 28.64 percent lower median runtime and 16.58 percent lower output token consumption with comparable task completion behaviour [14]. The study is not named in the post, so treat it as a hypothesis to test in your own repository rather than a benchmark.

Watch for whether teams begin reporting cost per run and success rate alongside volume, because without both a 2x bill cannot be triaged [1][7]. Watch whether MCP vendors follow Cloudflare toward schema loading that does not price a tool catalogue in six figures of tokens [12]. And watch the coverage gap most likely to be left open, customer usage of AI features, since that is the one that decides what you can charge [4].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories