Product1 publisher3 min readPublished
Claude Code's new /cost breakdown puts a token count on walking away from a session
Claude Code's /cost command now counts cache misses and, when it can, names a likely cause for the last one. Developers who leave sessions open over a break can now see how much usage a cold cache costs them.
The Product Desk · Product desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- A Prompt cache (main) line reports how much input came from cache, how many tokens must be cached again and whether the cache is currently warm or cold.
- The dollar estimate that /cost displays is not an additional charge for anyone on a Pro or Max subscription.
- Subscribers also see recent usage attributed to skills, subagents, plugins and individual MCP servers, viewable over the last day or the last week.
- An XDA Developers writer says the readout showed how many tokens their habit of leaving sessions open over breaks had been costing.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- decision A developer who keeps sessions open to avoid re-explaining a task can now set that saved context against the re-cache count each cold return produces.
- constraint Stopping points, named tests and a two-attempt rule limit spend only while Claude is running; an idle session that goes cold is diagnosed after the fact.
- capability Because subagents and MCP servers make their own requests, per-component attribution lets a subscriber tie usage to the specific plugin or subagent that drew it.
- cost For Pro and Max subscribers a cold cache is paid in usage, since re-cached context and cached reads both draw on it whatever the dollar estimate shows.
A developer comes back from a break to the Claude Code session they left open and asks it to continue. An XDA Developers writer who works this way says that request can cost more than anything sent before the break, because a long enough pause lets the prompt cache expire [7]. The writer keeps sessions open on purpose. The conversation still holds the files, instructions and decisions for the task, so nothing has to be explained again [5]. "For me, walking away is the biggest problem," the writer wrote [6].
Here's what teams tell themselves users do: a session that produced little code used little. Here's what this user actually does, by their own account: leaves the session open, walks off, and reaches for /cost in exactly the sessions where the amount of code Claude produced does not explain the usage [8]. Part of that gap is delegation. Subagents and other components make their own requests and draw usage alongside the main conversation [9].
The piece uses "walking away" for two different cases. In one, Claude is waiting. The cache expires, and the session's context has to be cached again before the next answer [7][1]. In the other, Claude is still working with nobody there to interrupt it. It may keep trying fixes that fail or search the same files again, and every attempt adds code and command output that goes out with each later request [10]. Caching makes that repeated input cheaper, but cached reads still consume usage [11].
Every fix the writer offers is for the working case. They name the affected feature and the test that should pass, and ask Claude to stop and explain if the same error survives two attempts [12]. For larger changes, plan mode puts the approach up for review before any edits, and the rewind command returns to an earlier checkpoint when an approach should be dropped [13]. The writer did not publish the token count the readout showed, or say how long a break has to last before the cache goes cold [14].
The useful way to sort any step-away is on two axes: whether Claude is running or waiting, and whether the absence is long enough for the cache to expire.
- Running, back soon: a loop can still be interrupted, and spend follows the work. - Running, gone long: the stopping point has to be set before leaving, as a named test and a two-attempt limit [12]. - Waiting, back soon: the cache should still be warm [7]. - Waiting, gone long: expect a cold cache. The first message back is the one to check against the Prompt cache line's re-cache count [3].
I'd keep the sessions open, since the saved context is the reason the writer works this way, and treat the re-cache count after a long break as what that context costs [5][3]. The tradeoff is that the re-cache comes back after every long break, so a developer who takes several a day pays it several times [7].
What to watch
- Whether Anthropic documents how long the Claude Code prompt cache lasts in an idle session, or lets users extend it.
- Whether the likely-cause field names idle expiry explicitly, so a cold-resume miss can be told apart from other causes.
- Whether per-component usage attribution reaches team or admin views beyond the individual subscriber's /cost screen.