Build1 publisher3 min readPublished
Catching runaway agent spend means pricing every trace span while the run is live
AWS Cost Anomaly Detection works from Cost Explorer data up to 24 hours old, so a dollar alarm on an agent fires after the money is spent. OpenTelemetry's GenAI spec has no cost attribute either, so teams must price each span from cache-split tokens and sum the trace tree.
The Engineer · Build desk

What happened
- A proposal to add cost attributes to OpenTelemetry's GenAI conventions, PR #443, has been open since 9 August 2026.
- OpenAI bills a cached input read at 0.1 times the uncached rate and a cache write at 1.25 times, with a 1,024-token minimum cacheable prefix.
- OpenAI's Agents API documentation says its usage counts are "not a final bill" and that missing usage "does not mean zero usage."
- Provider spend limits sit at organization or project scope, and nothing in the default stack attributes cost to a single tool call.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint Dashboards that multiply logged tokens by one input rate cannot be trusted for agent workloads, where cache state can move the price of identical input 12.5-fold.
- exposure A retry loop inside one tool call stays invisible to account-level limits and to the tool span itself until someone sums its child inference spans.
- cost Each team maintains its own per-model, per-date price table and its own cost attribute name, and pays for a rename if the open OpenTelemetry proposal settles on something else.
- decision Budget enforcement belongs in the agent runtime. AWS anomaly detection becomes a backstop for whatever the span sums missed.
OpenTelemetry's GenAI conventions describe their own token counters as "a proxy for cost approximation" [3]. Turning that proxy into a price takes three numbers per inference span. On Anthropic's API, `input_tokens` counts only the tokens after the last cache breakpoint, so the real input is `cache_read + cache_creation + input_tokens` [9]. Each bucket bills at its own rate: 1.25 times base input for a five-minute cache write, 2 times for a one-hour write, 0.1 times for a read [8]. A one-hour write costs 20 times a read of the same token [21]. According to The Agent Loop's post, the table those multipliers sit on changes per model and per date [20].
In a well-behaved agent, most input should be cheap reads. OpenAI's own multi-turn agent example reports cache hits above 90% [17]. Anthropic's documentation, as the post quotes it, names the easiest way to lose that: put a date at the front of the prompt and "you pay for a fresh cache write on every request and never get a read" [10]. It is an expensive way to tell a model what day it is. The token count is the same in both cases, so only a priced span shows the difference.
Input is where the volume sits. The post says Claude 4 Sonnet reportedly reached 100 billion tokens a day on OpenRouter in September 2025, with 99% of them input accumulated in trajectories [16]. It also cites a widely reported month-long run of about a hundred agent instances that billed $1,305,088.81 across 603 billion tokens and 7.6 million requests, figures the author calls un-audited [15]. That averages to about 79,000 tokens and 17 cents a request [22][23]. Braintrust's cost guide says an invoice at that size "can show that spending increased, but it cannot explain which customer, feature, prompt change, retry pattern, or agent run caused the increase" [14].
Attribution comes from the trace tree. The `execute_tool` span that wraps a tool call has no usage fields [4]. Its cost has to be the sum of the inference spans beneath it, and a run's cost is the sum over its whole tree. The post prescribes exactly that: price each span, roll it up from the child inference spans, and alarm at runtime [19]. A reader, @hannune, wrote in reply to the author's earlier post: "Tag cost per call is the one that took me too long to add" [18]. Until the spec defines a cost attribute [1], the name you pick is local convention, and renaming it later is part of the adoption cost.
I would build two guards into the rollup. A span that returns without usage enters the sum as unknown and gets flagged, following OpenAI's warning that missing usage is not zero [11]. And because the price table ages [20], the span sum is the runtime alarm while the invoice is the daily reconciliation.
Per-span cost is also what makes optimization measurable. The post cites AgentDiet's trajectory trimming cutting input tokens 39.9 to 59.7 percent and total cost 21.1 to 35.9 percent at equal performance [13]. Those ranges describe AgentDiet's workload. They carry over to yours only if your runs hold a similar share of context that can be dropped without hurting the task, and the post's point is that you cannot measure that share without per-span numbers [13].
What to watch
- Whether OpenTelemetry merges PR #443, and which attribute names and price source it standardizes for GenAI cost.
- Whether OpenAI or Anthropic add spend limits below organization and project scope, such as per run or per key.
- An audited breakdown of the reported $1.3 million month-long agent run, split by cache reads, cache writes and uncached input.