Skip to content

Build1 publisher3 min readPublished

Bedrock serves xAI's Grok 4.7 through the Converse, Responses and Chat Completions APIs

Amazon Bedrock now serves xAI's 500K-token-context Grok 4.7 through the Responses, Chat Completions and Converse APIs. Trying it from an existing client takes little code, though Artificial Analysis found its gains cost about twice the output tokens per task.

The Engineer · Build desk

Illustration accompanying Bedrock serves xAI's Grok 4.7 through the Converse, Responses and Chat Completions APIs

What happened

  • Grok 4.7 exposes four reasoning-effort levels on Bedrock: low, medium, high and xhigh.
  • Requests go to the bedrock-runtime endpoint and name a cross-Region inference profile instead of a bare model ID.
  • The model accepts text and image input and returns text.
  • xAI says Grok 4.7 is built on a new, larger base model, trained with a longer reinforcement learning run weighted toward tasks that take many hours.
  • xAI has begun giving selected cyber security partners invite-only access to Grok 4.7's red-team capabilities for defense research.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost A team running at xhigh, the effort Artificial Analysis tested, should budget roughly twice the output-token spend per task at any fixed per-token rate.
  • constraint With Grok 4.7 at xhigh and Grok 4.6 at per-measure effort levels, the published comparison does not isolate the model's gain from the setting's.
  • constraint Grok 4.7 is reached through cross-Region inference profiles, so teams with data-residency rules have to check which Regions a profile covers before sending regulated data.

For a team already on Bedrock, trying Grok 4.7 mostly means changing the model identifier in a client it already runs [4][5]. That makes a trial on the team's own work cheap to set up. The evidence for how capable the model is comes from two sources, and neither one was produced on that work.

The first is xAI. The AWS post lists the suites where xAI reports gains, among them CursorBench, DeepSWE, Terminal-Bench and the Harvey Legal Agent Benchmark [12]. For the scores themselves, it sends readers to xAI's announcement of September 21, 2026 [8]. The second is Artificial Analysis, which runs its own evaluations and does not rely on developer-reported figures [13].

Two details in the Artificial Analysis results decide whether they transfer. The largest coding gains were on agents run in xAI's own harness [14]. xAI also says it trained the model to natively understand its Grok Bot harness [11]. A team whose agent loop defines its own tools and stopping rules is testing a different system. The gains carry over only as far as that loop behaves like xAI's.

The other detail is effort. Grok 4.7 was measured at xhigh, while Grok 4.6 ran at whatever level Artificial Analysis had reported for each measure [15]. Unless those were Grok 4.6's highest settings, some of the gap comes from the setting.

Effort also sets the bill. Artificial Analysis found the gains come with roughly double the output tokens per task [16]. At any fixed per-token rate, output cost per task roughly doubles with them [1]. The post calls the model's theme endurance [22], and the token count agrees. "That's why it pays to set the effort level deliberately rather than inheriting the default," the AWS post says [17].

The experiment I would run first is small. It takes one task set from the team's backlog and runs it through the team's own harness at high and at xhigh [3]. It records pass rate and output tokens per task, next to the model already in production. In my context, if xhigh buys a few points for twice the tokens, high goes to production and xhigh is kept for the hardest tasks.

The same run can test self-verification. xAI says the model checks its own work more carefully and makes better use of its 500K context on long tasks [10]. The AWS post argues that a model which checks its output before continuing fails less catastrophically on long trajectories, where an early mistake otherwise compounds through every later step [18]. If that holds, there should be fewer traces where an error made early survives to the final answer. A team can count those in its own logs.

Security teams have one more claim to check. xAI says Grok 4.7 lets only a small fraction of risky dual-use prompts through while rarely blocking legitimate security work [20]. It also calls Grok 4.7 the strongest model it has tested on refusals and jailbreak resistance [19]. The over-refusal half is the part a security team can measure on its own prompts.

What to watch

  • Bedrock's per-token output price for Grok 4.7; with tokens per task roughly doubled, it sets the cost of running at xhigh.
  • An independent coding-agent evaluation of Grok 4.7 in a harness xAI did not build.
  • A Grok 4.6 against Grok 4.7 comparison at matched effort levels, to separate model gains from the xhigh setting.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories