Build1 publisher3 min readPublished
Anthropic meters AWS Claude usage in one-cent units, and charges 10% more to pin inference to the US
The platform pricing page converts token spend into Claude Consumption Units at $0.01 each for a single AWS line item, and adds a 1.1x multiplier for US-only inference on Claude 4.6 and later.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Claude Platform on AWS bills through AWS Marketplace using Claude Consumption Units (CCUs). Anthropic rates token usage in USD at standard per-model, per-feature rates, applies any negotiated discount, converts the result to CCUs at $0.01 per CCU, and reports the CCU quantity to AWS Marketplace hourly. The AWS bill shows a single CCU line item.
- For Claude 4.6 and later models, using inference_geo: "us" applies a 1.1x pricing multiplier.
- The pricing page's description of CCU billing covers rating in USD, application of negotiated discount, conversion at $0.01 per CCU, and hourly reporting to the marketplace; it does not describe how fractional CCUs are rounded.
- Claude in Microsoft Foundry bills through the Azure Marketplace using CCUs with the same process: rating in USD at standard rates, negotiated discount, conversion at $0.01 per CCU, hourly reporting, and a single CCU line item on the Azure bill.
- When you sign up on the AWS Console Claude Platform on AWS service page, the AWS Console looks up any private offer associated with your account and prompts you to accept it in AWS Marketplace. Anthropic directs customers to their account representative for private offer terms.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Anthropic's platform pricing documentation now describes two billing mechanics that a budget owner has to model separately. Usage on Claude Platform on AWS is rated in dollars, converted into Claude Consumption Units at $0.01 per CCU, and reported to AWS Marketplace hourly as one line item [1]. On Claude 4.6 and later, requesting US-only inference applies a 1.1x multiplier to token pricing [2]. One changes what finance can see; the other changes what data residency costs.
The order of operations on the AWS path matters. Anthropic rates token usage in USD at standard per-model, per-feature rates, applies any negotiated discount, converts the result to CCUs at a penny each, and reports the quantity hourly [1]. A dollar of net spend is therefore 100 CCUs [1]. Because the discount is applied before conversion and the invoice carries a single CCU line, the AWS bill itself will not attribute spend to a model or a feature; that reconciliation stays on your side, against your own token counts [2]. The page describes rating, discounting, conversion and hourly reporting, and says nothing about how fractional CCUs are treated [3]. If you meter chargeback per team, ask before you assume truncation or rounding is immaterial at your volume. Claude in Microsoft Foundry works the same way through Azure Marketplace [4]. Signing up on the AWS Console service page triggers a lookup of any private offer attached to your account, which you accept in AWS Marketplace; Anthropic points you to your account representative for the terms [5].
The geography multiplier is the more consequential number. Setting `inference_geo: "us"` on Claude 4.6 and later costs 1.1x across every token category: input, output, cache writes and cache reads [6]. Global routing remains the default at standard pricing [7]. It applies to the first-party Claude API and to Claude Platform on AWS, and on Claude in Microsoft Foundry the same 1.1x lands on deployments using the US Data Zone Standard deployment type [8]. Bedrock and Google Cloud, as partner-operated platforms, have independent regional pricing [9]. Earlier models do not accept the parameter at all and return a 400 error if you send it [10], which is a real constraint for anyone running one client library across a mixed model fleet.
The multipliers stack, including with the Batch API discount and data residency [11]. Prompt caching charges a cache read at 10% of the standard input price, with cache writes at 1.25x base input for the 5-minute duration and 2x for the 1-hour duration [12]. Under US-only inference, a cache read becomes 0.11x base input and a 1-hour write becomes 2.2x [3][4]. The break-even points are unchanged, because the same 1.1x sits on both sides of the ratio: one cache read for the 5-minute duration, two for the 1-hour duration [13][5]. Pinning geography raises the bill by a tenth; it does not change your caching architecture.
One asymmetry worth noting: fast mode, in research preview, offers faster output for Claude Opus 5 and Claude Opus 4.8 at premium pricing across the full context window, including requests over 200k input tokens, and is available on the first-party Claude API only [14].
Watch whether the CCU line item ever gains breakdown fields in AWS Cost Explorer, whether Anthropic publishes rounding behaviour for fractional CCUs, and whether the 1.1x US premium holds as more models ship past 4.6.