Build1 distinct publisher3 min readUpdated
A worked example on Anthropic's published August 2026 rates puts a five-agent fleet at $4,125 a month, nearly three quarters of it input. The post's own caching maths does not add up.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Take the fleet apart per call. $4,125 across 30,000 requests is 13.75 cents a request [1], and ten of those cents are spent before the model writes a character: 20,000 input tokens at $5 per million [2][2]. The generated output is 3.75 cents of it [2]. That is what the 73% figure means in practice [6], and it is why approval gates and per-team limits arrive after the money is gone.
Anthropic bills cache reads at a tenth of base input, with writes at 1.25x for a five-minute window or 2x for an hour [7]. The post's snippet declares 40% of input a stable prefix. Run the terms as written: 240 MTok at 0.1x is $120, the remaining 360 MTok at full rate is $1,800, output stays at $1,125, total $3,045 [3]. That is a 26% cut [4], not the roughly 9% the prose claims [8]. The number printed under the code, $3,765, is exactly what a 9% saving looks like [5], so the comments and the conclusion cannot both be right. The snippet also prices no cache writes at all [10], and writes are billed above base rate [7], so whatever the correct figure is, prefix churn pulls it down.
The routing number needs no reconciliation. Same fleet, same context, Haiku 4.5: $825 [10], which is 80% off and $3,300 a month back [6]. Divide the Opus bill by five and each agent costs $825 [7], so one Opus agent costs what the whole fleet costs on Haiku. Add the 50% batch discount that both major providers offer [11] on latency-tolerant traffic and the same work lands at $412.50 [8].
Price levers stop paying eventually, because retrieval payloads grow with document count: more documents match [12]. The post's example is four chunks totalling about 7,000 tokens, two of them contradicting each other, against 200 tokens of resolved facts carrying validity dates and provenance [12]. The contradiction is the part that costs money, since it is what makes an agent retry. On the vendor's own Terminal-Bench 2.1 run, a task-scoped memory layer used 41.2% fewer tokens at 72.6% lower cost, with mean reward moving from 83.37% to 88.31% across 445 trials [13], a gain of 4.94 points or 5.9% relative [9]. Cost fell further than tokens because fewer retries and shorter runs compound with smaller payloads [14]. The author says to read the methodology rather than trust the number [13], which is the right instruction to take literally.
The one step whose payoff does not depend on which of these figures is correct is the bucket split: system prompt, retrieved context, conversation history, output [15]. A day of work, and it tells you which of the four you have been arguing about for no reason.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Anthropic's published rates as of August 2026 for Opus 5 are $5 per million input tokens and $25 per million output tokens.
Haiku 4.5 is priced at $1 per million input tokens and $5 per million output tokens.
The worked example assumes 5 agents, 200 model calls each per day, 30 days, 20,000 input tokens and 1,500 output tokens per call: 30,000 requests, 600 MTok input and 45 MTok output per month.
On Opus 5 rates that fleet costs $4,125 a month: $3,000 of input plus $1,125 of output.
73% of the $4,125 monthly bill is input tokens, and the post argues most cost work targets the wrong 27%.
Anthropic bills cache reads at 0.1x the base input rate, with writes at 1.25x for a five-minute window or 2x for one hour, so a stable prefix is 90% cheaper to resend.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Arithmetic reproducible, impact claims not
The pricing-based core - rates, 30,000 requests, $4,125, the 73% input share, $825 on Haiku, half-price batch - recomputes exactly from figures published in the post, which is unusually checkable. Against that, the caching example contradicts its own code, the headline efficacy result is a self-run benchmark with no published methodology, and the framing anecdote is unsourced hearsay. Single publisher, no corroboration.
Pricing structures observable, deployment evidence minimal
Adoption signal is limited to published provider pricing and discount structures that any team can use, one second-hand enterprise usage disclosure, and one vendor-run benchmark. There is no independent deployment, customer or usage data for the compiled-facts knowledge layer the post advocates.
Mildly overstated, with one figure understated
Overstatement comes from the vendor framing: a self-run benchmark supplies the cost-down-accuracy-up claim, the 'tokenmaxxing versus contextmaxxing' thesis is asserted, quality risk from downgrading models is waved away, and the recommended remedy is the author's product. Understatement runs the other way in the caching section, where the printed ~9% saving is smaller than the ~26% its own terms imply. Net effect is modestly overstated relative to the evidence supplied.
Author sells the recommended remedy
The post discloses that the author works on Sentra, a company brain serving resolved cross-system facts to agents over MCP, which is explicitly the fourth and most-favoured lever in the article. It also routes readers to the vendor's own token calculator, statistics page and essay, and the only efficacy measurement is the vendor's own benchmark. Disclosure is present and explicit, which mitigates but does not remove the conflict.
Moderate on pricing maths, low on conclusions
Confidence is high that the pricing arithmetic and lever sizes are correct because they are recomputable from published rates, and low that the article's central conclusion holds, given one publisher, one vendor author, an internal arithmetic contradiction and an unverified benchmark.
build
Tier the models; the validation boundary is the thing you are actually buying1 distinct publisher
build
Claude's system prompt grew ninefold in two years. Version yours like code.1 distinct publisher
build
Agent reliability is a harness problem, not a prompt problem1 distinct publisher
build
Safety fixes ship in new model versions. The regression stays with whoever pinned the old one.1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 23, 2026