Build1 distinct publisher3 min readPublished
List is $5 per million tokens in and $25 out, with a 1M-token window. The figures that decide an agent budget come from multiplying those against turn counts and Anthropic's rate limit tiers.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
A single request that fills the 1M-token window costs $5.00 in input before the model writes a word [3][4][1]. Take the reply to the 128k ceiling and that adds $3.20 [2], so the most expensive single call claude-opus-5 will accept is $8.20 [3]. That ceiling is not the interesting number. The loop is.
Agents re-send their state every turn. A tool-using run that carries 500k tokens of context through ten turns bills 5M input tokens, which is $25 at list [4], and ten turns is unremarkable for a coding agent. The material we have quotes input, output, Fast mode and batch rates but no cache-read rate for this model [14], so the arithmetic above assumes you pay full price to re-read the same context on every turn. If caching applies, it moves the total more than any discount below.
Rate limits and budget are the same conversation. Anthropic's published tiers run from Tier 1 at 50 requests and 500k tokens per minute to Tier 4 at 400 requests and 4M tokens per minute, per documentation the post dates to May 2026 [9]. One full-context request therefore needs two minutes of a Tier 1 account's entire token allowance [5], which makes the 1M window a spec at that tier rather than something you can spend. At the top tier, saturating 4M input tokens a minute costs $20 a minute, or $1,200 an hour [6]. Committed-use discounts and dedicated throughput are negotiable, with enterprise contract approval running 3 to 10 business days against instant self-serve activation [10].
Fast mode is priced at 2x for roughly 2.5x throughput [5], which works out to $10 and $50 per million [9]. It does not reduce the token count, so the trade is double the spend for about 40% of the wall clock [7]. Defensible for interactive coding, hard to justify overnight, particularly against the 50% discount the documentation describes for asynchronous batch work, whose exact Opus 5 rate the post says to confirm on the official pricing page [6].
The gateway route is the one number that moves unit economics. Api.Airforce lists $3.58 and $17.88 per million through an OpenAI-compatible endpoint, roughly 28% under list [7]. On the ten-turn run above, that is $17.90 of input instead of $25, a saving of $7.10 [8]. The same post warns that gateway pricing is volatile [8].
Worth noting where all this comes from: a dev.to writeup that attributes the rates to Anthropic's pricing page and newsroom alongside several aggregator sites, and that repeatedly tells readers to check the official page rather than trust the article [13]. The $5 and $25 are firm enough to plan against. The batch and cache rates are not numbers yet.
Ranked by verification strength, evidence, and original report placement.
The dev.to post attributes its figures to Anthropic's pricing page, model documentation and newsroom plus aggregator sites, and repeatedly instructs readers to check the official pricing page or model documentation.
The source lists input, output, Fast mode and Batch API rates but states no prompt-caching or cache-read rate for claude-opus-5.
Filling the full 1M-token context window once costs $5.00 in input tokens at list price.
A response at the 128k output ceiling costs $3.20 at list price.
The most expensive single call the model will accept, full context in and full output out, is $8.20 at list price.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single secondhand publisher, self-hedged figures
Every external fact rests on one dev.to post that routes its numbers through aggregator sites (moclaw.ai, wavespeed.ai, layer3labs.io, bleap.finance, api.airforce) and repeatedly tells readers to confirm on Anthropic's official pricing page or model documentation. Some inputs are dated May 2026 rather than the article's Aug 2026 framing, and the batch rate is explicitly left unconfirmed. The only internally verifiable material is the article's own sourcing behaviour, the volatility caveat, the absence of a caching rate, and arithmetic derived from the figures it states.
Availability breadth claimed, no usage evidence
The cluster documents distribution surface only: a release date, six access channels including Bedrock, Vertex AI and GitHub Copilot tiers, and a reseller gateway listing. There is no usage disclosure, customer count, traffic figure, deployment case or benchmark result anywhere in the source, and the availability claims themselves are secondhand. Availability is scored as a weak adoption signal rather than inferred uptake.
Rate card presented with more certainty than its sourcing supports
The post presents precise prices, specs, tier limits and a 28% gateway undercut in the register of settled fact while simultaneously hedging nearly every figure, omitting the caching rate that dominates agent-loop cost, and recommending gateways whose pricing it calls volatile. The concrete budget consequences of the numbers, per-call ceilings, ten-turn loop cost, and the $1,200-per-hour Tier 4 saturation figure, are absent from the coverage and only appear as derivations, so the practical cost picture is overstated in favour of the headline rate.
Gateway resale promotion inside a developer explainer
The article repeatedly steers readers to unified gateways, naming HeFu as an access path and listing its wider model catalogue, and headlines a third-party gateway's own quoted rates as a 28% discount on Anthropic list. That is commercial promotion of resale channels embedded in what reads as neutral setup documentation, and the gateway rates are sourced from the gateway. The main mitigation is that the post does publish volatility and feature-coverage warnings alongside the pitch.
Low: one hedged source, arithmetic sound but inputs unverified
Confidence is limited by a single-publisher cluster with no primary documentation, aggregator-chained attribution, mixed as-of dates, and no independent corroboration of pricing, specs or tier limits. What can be held with more confidence is narrow: the article's own sourcing and disclosure behaviour, and the arithmetic that follows from the figures it publishes, which is internally consistent even if the inputs are not confirmed.
build
Claude Code now opens in auto mode: a classifier, not you, approves the shell commands1 distinct publisher
product
Anthropic's usage policy says no explicit content. Opus 4.6 said yes 10 times out of 10.1 distinct publisher
product
Four leaderboards, four denominators: what you buy when you standardize on a coding agent1 distinct publisher
build
Opus 5 absorbed your verify prompts. The reading is still on your desk.1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 27, 2026