Build1 publisher3 min readPublished
Six token trackers, no referee: LLM spend dashboards disagree by 2x to 8x
An independent recount of six usage trackers found gaps of 2.00x to 8.09x. The bigger problem is that none of them, and no invoice, can say which number was right.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Over five months the author took apart the folding logic in six token usage trackers and recomputed each one with an independent implementation.
- The trackers do not agree; the gaps run from 2.00x to 8.09x, in both directions.
- Ten fixes have landed so far across five repositories.
- None of the audited tools knows whether its own number is right; there is nothing for them to check against.
- tokscale issue #1011 was filed by the author in July 2026 and was still open and unlabelled as of Aug 18.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A developer publishing on dev.to spent five months pulling apart the folding logic in six token usage trackers and recomputing each one with an independent implementation [1]. The totals disagree by 2.00x to 8.09x, in both directions [2], and the second finding is the one that should worry anyone running a cost or context-budget dashboard: none of these tools knows whether its own number is right, because there is nothing to check it against [4]. The audit did not begin as a bug hunt. According to the author, they were reading their own session logs with two tools and noticed the totals did not match; a rounding-sized gap would have been written off, but this was close to double [7]. The failure modes are structural rather than sloppy. First, streaming double counting: a streaming response emits a usage chunk and then assembles a complete message when the request finishes, both carrying the same usage, so summing everything you see yields exactly 2x [11]. On the deepseek-harness corpus, naive folding came out at 2.000000x of the official projection, with two provider routes computed separately agreeing to six decimal places [12]. The same hole wears different clothes as per-content-block counting: claude-code-templates, at 30,276 stars, measured 2.36x [13], and Clawdmeter measured 2.34x to 2.37x before its author shipped v3.0.1 on 2026-08-08 under the release title "token counts corrected (~2.5x lower)" [14]. That release also noted that the percentage progress bar was never affected, because it reads Anthropic's rate-limit header rather than the folded total [15]. Second, session forks. A child session's log file physically contains the parent's full prefix, so any aggregation that walks every session file and sums counts that prefix twice [16]. The two forks the author caught were inflated 5.18x and 23.05x, and the ratio gets worse the later the fork point, because the inherited prefix grows while the child's own contribution shrinks [17]. The correct test is seq >= seedLength; checking origin === 'subagent' alone misses ordinary user-created forks, which inherit a prefix too [18]. Third, compaction, which fails in the other direction. Context compaction is a real model call and the provider reports usage for it, but it is not a loop step, so it produces neither an assistant chunk nor an assistant message, and folding logic that matches only those two event types never sees it [19]. Measured: three compaction events, 48,895 tokens, none counted [20], an average of roughly 16,300 tokens per event missing from the ledger [23]. The largest summarize call reported 44,444 tokens, 41,472 of them cache reads, to replace a history range of 19,962 tokens [21], about 2.2x the tokens it removed [24]. Compaction fires most often on long sessions, which are exactly the sessions where the number matters [22]. None of this would be more than a patch queue if there were an invoice to reconcile against. For subscription users there is not: the fee is flat, so Anthropic never sends a per-token bill [9], and Claude Code's own dollar figure carries a docs warning that usage for Max and Pro subscribers is included in the subscription and so is not relevant for billing purposes [10]. Where a gap is large enough to see, it does close fast. In tokscale issue #76, johnpyp reported a near-2x difference against ccusage; maintainer junhoyeo labelled it a bug within twenty minutes and closed it within five hours, tracing the cause to Claude Code upstream duplicating session history under stream-json rather than to tokscale [8]. Ten fixes have landed across five repositories so far [3]. The author is upfront that two still-open items are their own filings: tokscale issue #1011, filed July 2026 and still open and unlabelled as of Aug 18 [5], and claude-code-templates PR #754, still unmerged as of Aug 18 [6]. That is a fair disclosure and it lowers the temperature: "not fixed" here means "not yet picked up", not disputed.