Build1 publisherNot yet confirmed elsewhere3 min readPublished
Bedrock's two GPT-6 Astra endpoints split prompt caching from invocation logging
Amazon Bedrock serves GPT-6 Astra through two endpoints that split explicit prompt caching from invocation logging, a dev.to analysis of the model card finds. Teams have to choose per request path, and the model's end-of-life, no sooner than September 8, 2027, sets how long that choice must hold.
The Engineer · Build desk

What happened
- On bedrock-runtime the model answers Converse, Responses and Chat Completions, but it does not answer InvokeModel.
- Above 272,000 input tokens the whole request reprices, to $20 input and $75 output per million tokens on the global profile.
- GPT-6 Astra reached general availability on Bedrock on September 8, 2026, with a 1,050,000-token context window.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision A workload that needs Guardrails or content-level logging gives up the cache rate and pays ten times as much for every repeated prefix it sends.
- exposure Incident reviews and compliance evidence for mantle traffic depend entirely on what the application logs itself, because Bedrock keeps only caller records for that path.
- cost Without CountTokens, a prompt that drifts past 272,000 tokens more than doubles its input bill unnoticed, so teams need their own pre-flight count with a margin.
- constraint Shared clients built on InvokeModel cannot reach Astra, so a single instrumentation point has to be rebuilt around Converse or the Responses API before adoption.
The split follows the control plane. Guardrails exist only on bedrock-runtime, and only through Converse [7]. Runtime also serves the Responses and Chat Completions APIs [5]. Those calls are synchronous only, and a request with background set returns a 400 [8]. Converse and Responses calls on runtime are captured by invocation logging to S3 or CloudWatch, with 100 KB inline [9]. bedrock-mantle exposes the OpenAI-shaped paths, /openai/v1/responses and /openai/v1/chat/completions [6], with async calls and server-side tools [10].
The cache is the well-built part. Breakpoints are explicit: a request sets up to four with `prompt_cache_breakpoint`, each cached prefix needs at least 1,024 tokens, and entries expire after 30 minutes [11]. A reviewer can see in the code which prefix is reused and for how long. Reads cost $1.10 per million tokens [12], one-tenth of the $11 in-Region input rate [23].
That discount is paid for in audit coverage. Mantle requests never reach the invocation log. CloudTrail records the bedrock-mantle:CreateInference action and its caller, without the content [13]. Separately, Responses calls default to `store=true`, keeping input and output for 30 days, and a call routed through the global profile is stored in the Region that served it [14]. The global profile can route to any commercial Region [16].
Existing clients break first. A wrapper standardized on InvokeModel to get one instrumentation point does not call Astra [5]. CountTokens is marked unsupported on bedrock-runtime for this model, as are structured outputs [20]. Above 272,000 input tokens the whole request reprices, not just the overflow [18]. A 280,000-token prompt on global costs $5.60 in input, and the same prompt trimmed to 270,000 costs $2.70 [19]. The prompt grew 3.7% and the input bill grew 107% [24]. Bedrock drew a price line at 272,000 tokens and left out the API that counts them, so any pre-flight check has to run on the client [20].
Capacity has one setting. Standard is the only supported tier. Priority, Flex and Reserved are not, so Region is the only control over tail latency [15]. The us. profile covers five Regions in the US and Canada [16]. Geographic and in-Region routing cost $11 per million input tokens against $10 on global [17], a 10% premium for knowing where inference runs [26].
In my context, a regulated workload where Guardrails are mandatory, I would put the audited path on runtime Converse with a geographic profile and pay the uncached $11 per million input tokens [17]. A cache-heavy internal agent goes to mantle, with request and response capture built into the application, because no Bedrock log will hold that content [13]. The lifecycle sets how long that split has to hold: end-of-life no sooner than September 8, 2027, with a legacy period of at least six months [21]. GA to the earliest end-of-life is twelve months [25]. The author of the dev.to analysis wrote that this is "a one-year contractual floor for planning migration, and it is the kind of fact that belongs in your ADR, not in your slide" [22].
What to watch
- Whether AWS adds invocation logging to bedrock-mantle or explicit caching to bedrock-runtime, which would remove the per-path choice.
- Whether CountTokens and structured outputs become supported for openai.gpt-6-astra on bedrock-runtime.
- Whether Priority, Flex or Reserved tiers open for Astra, giving a tail-latency control beyond Region choice.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap0
- Incentives
- Insufficient
- Confidence45
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
GPT-6 Astra reached general availability on Amazon Bedrock on September 8, 2026, with a 1,050,000-token context window.
- [2]
openai.gpt-6-astra has a 128,000-token output ceiling and an April 30, 2026 knowledge cutoff.
- [3]
Astra accepts text and image input; output is text only.
- [4]
The model is served by two distinct endpoints, bedrock-runtime and bedrock-mantle, which do not offer the same guarantees; the model card does not provide explicit prompt caching and invocation logging on the same path.
- [5]
On bedrock-runtime the model answers on Converse, Responses and Chat Completions, and not on InvokeModel; an internal wrapper standardized on InvokeModel for a single instrumentation point does not call Astra.
- [6]
On bedrock-mantle only Responses and Chat Completions are available, served at /openai/v1/responses and /openai/v1/chat/completions.
- [7]
Guardrails are available only through the Converse API on bedrock-runtime.
- [8]
On bedrock-runtime, Responses and Chat Completions are sync only; a background request returns a 400.
- [9]
Invocation logging to S3 or CloudWatch, with 100 KB inline, captures request and response for Converse and Responses calls on bedrock-runtime.
- [10]
bedrock-mantle offers the Responses /v1 API with server-side tools and async calls.
- [11]
The explicit prompt cache on bedrock-mantle uses prompt_cache_breakpoint, requires a minimum of 1,024 tokens, allows 4 breakpoints and has a 30-minute TTL.
- [12]
bedrock-mantle serves Astra in-Region in us-west-2, and cache reads cost $1.10 per 1M tokens.
- [13]
Requests on bedrock-mantle fall outside the invocation log; CloudTrail records the bedrock-mantle:CreateInference action and who called, not the content.
- [14]
store=true is the default and stores input and output for 30 days; a call on the global profile is stored in the Region that served it.
- [15]
Astra supports only the Standard service tier; Priority, Flex and Reserved are listed as unsupported, leaving Region choice as the only control over tail latency.
- [16]
The us.openai.gpt-6-astra profile covers 5 Regions in the US and Canada; global.openai.gpt-6-astra can route to any commercial Region.
- [17]
Global cross-Region inference costs $10.00 per 1M input tokens and $50.00 per 1M output; in-Region and Geo cross-Region inference cost $11.00 and $55.00.
- [18]
Above 272,000 input tokens the entire request reprices: $20.00 input and $75.00 output per 1M on global, $22.00 and $82.50 in-Region.
- [19]
A 280,000-token prompt on global costs $5.60 of input; the same prompt trimmed to 270,000 tokens costs $2.70.
- [20]
The bedrock-runtime feature list marks CountTokens and structured outputs as unsupported for this model.
- [21]
Astra's end-of-life is no sooner than September 8, 2027, with a legacy period of at least 6 months.
- [22]
"that is a one-year contractual floor for planning migration, and it is the kind of fact that belongs in your ADR, not in your slide."
- [23]
Cache reads on bedrock-mantle cost one-tenth of the in-Region uncached input rate.
- [24]
Growing a global-profile prompt from 270,000 to 280,000 tokens adds 3.7% to its size and 107% to its input cost.
- [25]
The time from GA to the earliest end-of-life date is twelve months.
- [26]
In-Region and Geo routing carry a 10% premium over global on input price.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toGPT-6 Astra on Bedrock: the endpoint decides what you audit
1 article · October 8, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- LLM Inference PricingFollow
- Prompt CachingFollow
- Managed model hostingFollow
- AI audit loggingFollow