Skip to content

Build1 publisher3 min readPublished

Bedrock turns GPT-5.6 throughput into a routing choice, with residency as the price

Cross-Region inference for GPT-5.6 on Amazon Bedrock means capacity ceilings are now fixed by changing a profile prefix, not by changing models. The tradeoff is where your data gets processed.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Amazon Bedrock now offers OpenAI GPT-5.6 models in more than 25 AWS Regions, with cross-Region inference.
  • Three GPT-5.6 variants support cross-Region inference: Sol, Terra and Luna, each tuned for a different balance of capability and cost. They are the general-purpose variants; the family also includes specialized cyber security variants.
  • Cross-Region inference (CRIS) is primarily a capacity mechanism: by allowing requests to draw on a broader pool of compute rather than being bound to one Region's available capacity, it improves throughput and helps maintain consistent performance under load.
  • Amazon Bedrock inference profiles are logical identifiers you pass instead of a raw model ID.
  • An inference profile defines a model and the AWS Regions to which Amazon Bedrock can route a request.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Amazon Bedrock has switched on cross-Region inference for OpenAI's GPT-5.6 models, which it says are now available in more than 25 AWS Regions, with three general-purpose variants supporting the feature: Sol, Terra and Luna [1][2]. According to AWS, cross-Region inference is primarily a capacity mechanism, letting requests draw on a broader pool of compute instead of being bound to one Region's available capacity, which improves throughput and helps hold performance steady under load [3]. That relocates the usual fix for a throughput ceiling from the model catalogue to the identifier you pass at call time.

The mechanics are deliberately boring, which is the point. An inference profile is a logical identifier you use instead of a raw model ID, and it defines both a model and the set of Regions Bedrock may route to [4][5]. You call the profile from a source Region, and Bedrock routes the request to a destination Region and runs it on compute there [6]. A geographic profile carries a geography prefix, for example us.openai.gpt-5.6-terra, and can only route to destinations inside that geography [7]. A global profile, prefixed global. as in global.openai.gpt-5.6-terra, can route to any supported commercial AWS Region where the model is deployed, chosen on real-time capacity [8]. For GPT-5.6 this launch ships exactly two options: US geographic and global [9], which across three variants is six routing identifiers to choose between [14].

Two details make this a tuning knob rather than a migration. Billing and quota consumption are tracked against your account regardless of which backend Region handled the request, so you keep one spending and throughput picture [10]. And the three variants are otherwise interchangeable on capability surface: all accept text and image input and return text, all have a 1 million token context window, and all support reasoning mode, server-side tool calling and prompt caching [11], callable through the OpenAI Responses API, the OpenAI Chat Completions API and the Bedrock Converse API [12]. AWS describes the variants as tuned for different balances of capability and cost [2]. So model selection is a cost and quality decision; the prefix is the capacity decision.

The cost of the widest capacity pool is stated plainly by AWS: data processed through global cross-Region inference may cross the Regions in that model's eligible set, and if your workload has residency requirements restricting processing to specific geographies, you should use the geographic profile for that geography or call a single Region directly [13]. Since the only geographic profile at launch is US, a team with non-US residency constraints is left with direct single-Region calls, which is the configuration cross-Region inference exists to escape [15]. That is the real asymmetry here: US workloads get an elastic capacity story, everyone else gets a pinned one.

Two things to check before you rewrite a model ID in production. First, which Regions actually participate in each profile set: AWS points to its cross-Region inference documentation for the per-model global and geographic lists [16], and the launch post's own Region tables are the authority on source versus destination, with the same routing applying to all three variants [17]. Second, prompt caching is listed as supported on all three variants [11], but the post does not describe how cache behaviour holds up when consecutive requests land in different destination Regions. If your unit economics depend on cache hits, measure that before you move traffic to a global profile.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories