Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished

SpaceX's Grok Bot will route work to Claude Opus 5.5 whenever Anthropic's model does it better

Elon Musk said SpaceX's Grok Bot agent will send each task to the best back-end model, Anthropic's Claude Opus 5.5 included. The policy moves the engineering work into routing, and into the tests that decide what 'best' means for each task.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying SpaceX's Grok Bot will route work to Claude Opus 5.5 whenever Anthropic's model does it better
Generated illustration

What happened

  • Grok Bot is a joint product of SpaceXAI and Cursor, which SpaceX agreed in June to buy for $60 billion in stock in a deal that closed in August.
  • At least 10 of 22 company leaders interviewed at Insight Partners' ScaleUp:AI conference said they pick the model by task and save the expensive one for work that needs it.
  • In the week through October 7, four of OpenRouter's ten most-used models were cheap, fast "Flash" tiers that labs ship alongside their flagships.

Why it matters

  • cost Every model added to a routing table needs its own regression tests, so part of the token saving goes back out as eval engineering.
  • exposure Grok Bot's output now depends on release schedules at Anthropic, MidJourney and Suno, so a vendor's pricing change that alters model behaviour reaches SpaceX's users.
  • decision Anyone evaluating Grok Bot has to ask which back end handled a given task, since the same agent may answer with Grok or with Claude.

Musk's Wednesday note sets a policy [4]. "Whatever is most likely to give you the best outcome," he wrote [5]. That phrase is a forecast made per task, before the task runs. Some component has to make it for every request Grok Bot receives. The post does not say what scores the candidate models, or against which tests [5].

The operators Matt Burns interviewed for The New Stack described what that component usually looks like [1]. The most common setup was a funnel [10]. One security company runs a rules engine over everything first, passes what survives to small models, and lets larger models see only what is left [10]. One of its executives said a petabyte of data would cost millions of dollars to process with any model, small ones included [11]. A second company sends each agent request through a cheap model to work out what the user wants before the expensive model does anything [12].

The classifier stage keeps getting cheaper. Haiku 5.5 now costs $0.10 per million input tokens after a 90% cut. That puts the earlier rate for requests under 100,000 tokens at about $1 [18]. Burns notes that small models like Haiku typically handle classification and routing [16].

Switching is where routing breaks. One engineering leader told Burns that three days before demoing to an important prospect, his team moved to a newer model that the benchmarks showed was cheaper and no worse [13]. It broke the workflow, and the team spent a weekend rolling back [13]. His lesson was evals [13]. Vendors can cause the same failure from their side. The New Stack's Amanda Caswell reported that when Anthropic made Opus 5.5 20% cheaper than Opus 5, the change broke four things agents depend on [14]. Burns's rule is "Always Be Switching" [17]. The demo team would want the rule to name a weekday.

Burns wrote in August that "when models converge on price, the money moves to whoever decides which model gets the job" [7]. His evidence supports a narrower claim. Cost was the reason most interviewees gave for routing [2]. The interviews count teams that route. That is evidence of adoption, while Burns's claim is about where the money goes. The sample is also narrow. The 22 companies are Insight Partners portfolio companies, interviewed for an Insight video series, and Insight owns The New Stack [8]. A couple of the interviewees sell small models for a living, and Burns says he weighs their enthusiasm accordingly [9].

We think the routing logic is the easy part to copy. A rules engine and a cheap intent classifier are ordinary components [10][12]. The harder part is the eval suite that defines "best" for each task and catches a vendor's breaking change before a customer does. If SpaceX gets a lasting advantage from Musk's policy, we'd expect it to come from those tests. A suite like that is what the demo team lacked [13].

What to watch

  • Whether SpaceX documents how Grok Bot chooses a back end per task, or shows users which model handled a request.
  • How Grok Bot behaves the next time Anthropic changes Opus pricing or behaviour that agents depend on.
  • Whether Flash-tier models hold four of OpenRouter's top ten slots once Haiku 5.5's lower price takes effect.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence50
Adoption55
Hype gap+25
Incentives70
Confidence50
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    At least 10 of the 22 company leaders interviewed at Insight Partners's ScaleUp:AI conference described picking the model by task and keeping the expensive one for the work that needs it.

    ReportedSupportedSource: Matt Burns, The New Stack2 sources— create a free account to open themView cited source
  2. [2]

    Cost was the reason most of the interviewed leaders gave for model triage.

    ReportedSupportedSource: Matt Burns, The New Stack2 sources— create a free account to open themView cited source
  3. [3]

    In the week through October 7, four of OpenRouter's ten most-used models were "Flash" models, the cheap, fast versions labs ship alongside their flagships.

    ReportedSupportedSource: OpenRouter leaderboard, via The New Stack2 sources— create a free account to open themView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. thenewstack.io

    1 article · October 10, 2026

    Elon Musk’s Grok Bot will pick Claude over Grok when it’s better. One-model loyalty is dead.

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories