Build1 distinct publisher2 min readPublished
Anthropic meters Pro and Max on a rolling 5-hour window and a weekly cap, both counted in tokens. What drains them is re-read context, not how many questions you ask.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The billing mechanism is unglamorous, and it is the whole story. Claude Code sends the full conversation with every request, and each time Claude uses a tool it fires another request carrying that batch of tool results, which the dev.to write-up attributes to Anthropic's own cost documentation [6]. The price of a message is set by what is stacked behind it, not by what you typed.
The post's worked example is worth keeping. Developer A opens a fresh session and asks 40 short questions about one small file. Developer B keeps a single session open from 9am across three unrelated tasks and sends 12 messages [7]. B burns more of the plan, because that one-line question at 4pm draws usage for every file read and every tool result sitting behind it [8][9].
Prompt caching is what normally keeps a long session survivable, since re-read history bills at the cached rate rather than the full one [10]. On a subscription that cache lives one hour [11]. The session window runs five [3], so a single window spans five cache lifetimes [1], and the first message after a long lunch reprocesses the entire context at full price [11].
Two defaults sit underneath all of it. Extended thinking is on by default, thinking tokens bill as output tokens, and the default budget can run to tens of thousands of tokens per request depending on the model [12]. Agent teams, when switched on, use roughly seven times the tokens of a normal session [13], which leaves about 14% of the run you would otherwise get from the same window [2].
Put that beside the price list. Pro is 20USD a month billed monthly, Max starts at 100USD, and both get the same model access [14]; Max is sold as 5 or 20 times Pro's usage [16]. Entry Max is therefore exactly five times the monthly Pro price for an advertised five times the usage [3]. Money buys headroom at par and buys no efficiency, and agent teams alone would swallow the upgrade. Clearing context buys the same headroom for nothing.
The weekly meter is the one that ends the week. It resets at a fixed hour assigned to the account, and you can stay inside every 5-hour window and still be finished on a Thursday [4]. The housekeeping that rescues an afternoon compounds over seven days.
One diagnostic decides whether housekeeping helps at all. A session or weekly limit message is plan-wide, and /model will not rescue you; an Opus or Sonnet limit is model-specific, and switching family genuinely keeps you working [17].
Ranked by verification strength, evidence, and original report placement.
Claude Pro meters usage on two clocks at once, a rolling 5-hour session window and a weekly cap; neither is a message counter and both are token meters. The 5-hour window starts with the first message and closes five hours later whether one message or four hundred were sent.
The weekly window resets at a fixed time assigned to the account, and a user can stay inside every 5-hour window all week and still run out on a Thursday.
Anthropic's documentation is explicit that usage across claude.ai, Claude Code and Claude Desktop counts toward the same pool.
Per Anthropic's cost documentation, Claude Code sends the full conversation with every request, and each time Claude uses tools it sends another request carrying that batch of tool results.
The post's example: developer A opens a fresh Claude Code session, asks 40 short questions about a small file and closes it; developer B opens one session at 9am, works across three unrelated tasks without clearing, and sends 12 messages.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single practitioner explainer; doc-grounded mechanics, unlinked citations
All claims trace to one dev.to post. The metering mechanics (dual windows, shared pool across surfaces, full-context resend per request, one-hour subscription cache, plan prices, limit-message scope) are attributed to Anthropic documentation and are internally consistent, which supports them at a working level. But no primary link, doc version or retrieval date is supplied, no second publisher corroborates, and the sharpest quantities — the roughly 7x agent-team draw, the 7% affected-user figure and the March 2026 anecdote — carry no measurement or citation at all.
Live production metering and pricing; no usage magnitudes disclosed
The subject is an in-production billing and rate-limiting regime, not a proposal: prices for Pro and Max are published, the metering spans claude.ai, Claude Code and Desktop, and a peak-hour limit adjustment was reportedly rolled out to paying subscribers with a stated 7% impact scope. Adoption is therefore real but only coarsely measurable here — no subscriber counts, no token quotas, no independent telemetry, and the one impact figure is an unverified vendor paraphrase.
Mildly overstated: firm numbers on soft evidence
The core thesis — limits are token meters and re-read context drains them — is modest and well matched to the documented mechanics, so the gap is small. It tilts positive because the post states unmeasured multipliers with confident precision (roughly 7x for agent teams, 'not even close' for the two-developer comparison, a 7% affected-user share), generalises an anecdote about a 21%-to-100% jump, and carries an author aside promoting their own project with a self-reported 40% overhead reduction that no cited benchmark backs.
Author promotes own adjacent project inside the explainer
The author discloses building ANRL, a Rust-based AI-native representation language aimed at delimiter overhead and context fragmentation, and cites a self-reported 40%+ reduction in delimiter token overhead from their own benchmarks. That gives a direct interest in framing subscription limits as a context-waste problem. The piece is published on a developer community platform where engagement rewards how-to framing. There is no evidence of vendor sponsorship, and the post argues against inflating unknown quotas, which moderates the score.
Low-moderate: consistent mechanics, one unverified publisher
Confidence is limited by structure more than content: a single publisher, no primary vendor source in the cluster, and no independent measurement. The documented mechanics hang together and are checkable in principle, which supports directional confidence about how the meters work; the specific multipliers, affected-user share and anecdote should be treated as unconfirmed.
product
A 2x LLM bill is not a bug report: token spend is an observability problem1 distinct publisher
build
A Retention Policy for Agent Memory: Flag Unused Skills at 30 Days, Archive at 901 distinct publisher
build
A session that read "finished" and "still executing" was a slow queue, not a dropped handshake1 distinct publisher
build
52 days of zeros: what a cost hook records when the payload never had the numbers1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 25, 2026