BuildNot yet confirmed elsewhere1 publisher3 min readPublished
SpaceX's Grok Bot will route work to Claude Opus 5.5 whenever Anthropic's model does it better
Elon Musk said SpaceX's Grok Bot agent will send each task to the best back-end model, Anthropic's Claude Opus 5.5 included. The policy moves the engineering work into routing, and into the tests that decide what 'best' means for each task.
The Engineer · Build desk

What happened
- Grok Bot is a joint product of SpaceXAI and Cursor, which SpaceX agreed in June to buy for $60 billion in stock in a deal that closed in August.
- At least 10 of 22 company leaders interviewed at Insight Partners' ScaleUp:AI conference said they pick the model by task and save the expensive one for work that needs it.
- In the week through October 7, four of OpenRouter's ten most-used models were cheap, fast "Flash" tiers that labs ship alongside their flagships.
Why it matters
- cost Every model added to a routing table needs its own regression tests, so part of the token saving goes back out as eval engineering.
- exposure Grok Bot's output now depends on release schedules at Anthropic, MidJourney and Suno, so a vendor's pricing change that alters model behaviour reaches SpaceX's users.
- decision Anyone evaluating Grok Bot has to ask which back end handled a given task, since the same agent may answer with Grok or with Claude.
Musk's Wednesday note sets a policy [4]. "Whatever is most likely to give you the best outcome," he wrote [5]. That phrase is a forecast made per task, before the task runs. Some component has to make it for every request Grok Bot receives. The post does not say what scores the candidate models, or against which tests [5].
The operators Matt Burns interviewed for The New Stack described what that component usually looks like [1]. The most common setup was a funnel [10]. One security company runs a rules engine over everything first, passes what survives to small models, and lets larger models see only what is left [10]. One of its executives said a petabyte of data would cost millions of dollars to process with any model, small ones included [11]. A second company sends each agent request through a cheap model to work out what the user wants before the expensive model does anything [12].
The classifier stage keeps getting cheaper. Haiku 5.5 now costs $0.10 per million input tokens after a 90% cut. That puts the earlier rate for requests under 100,000 tokens at about $1 [18]. Burns notes that small models like Haiku typically handle classification and routing [16].
Switching is where routing breaks. One engineering leader told Burns that three days before demoing to an important prospect, his team moved to a newer model that the benchmarks showed was cheaper and no worse [13]. It broke the workflow, and the team spent a weekend rolling back [13]. His lesson was evals [13]. Vendors can cause the same failure from their side. The New Stack's Amanda Caswell reported that when Anthropic made Opus 5.5 20% cheaper than Opus 5, the change broke four things agents depend on [14]. Burns's rule is "Always Be Switching" [17]. The demo team would want the rule to name a weekday.
Burns wrote in August that "when models converge on price, the money moves to whoever decides which model gets the job" [7]. His evidence supports a narrower claim. Cost was the reason most interviewees gave for routing [2]. The interviews count teams that route. That is evidence of adoption, while Burns's claim is about where the money goes. The sample is also narrow. The 22 companies are Insight Partners portfolio companies, interviewed for an Insight video series, and Insight owns The New Stack [8]. A couple of the interviewees sell small models for a living, and Burns says he weighs their enthusiasm accordingly [9].
We think the routing logic is the easy part to copy. A rules engine and a cheap intent classifier are ordinary components [10][12]. The harder part is the eval suite that defines "best" for each task and catches a vendor's breaking change before a customer does. If SpaceX gets a lasting advantage from Musk's policy, we'd expect it to come from those tests. A suite like that is what the demo team lacked [13].
What to watch
- Whether SpaceX documents how Grok Bot chooses a back end per task, or shows users which model handled a request.
- How Grok Bot behaves the next time Anthropic changes Opus pricing or behaviour that agents depend on.
- Whether Flash-tier models hold four of OpenRouter's top ten slots once Haiku 5.5's lower price takes effect.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence50
- Adoption55
- Hype gap+25
- Incentives70
- Confidence50
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
At least 10 of the 22 company leaders interviewed at Insight Partners's ScaleUp:AI conference described picking the model by task and keeping the expensive one for the work that needs it.
ReportedSupportedSource: Matt Burns, The New Stack2 sources— create a free account to open themView cited source - [2]
Cost was the reason most of the interviewed leaders gave for model triage.
ReportedSupportedSource: Matt Burns, The New Stack2 sources— create a free account to open themView cited source - [3]
In the week through October 7, four of OpenRouter's ten most-used models were "Flash" models, the cheap, fast versions labs ship alongside their flagships.
ReportedSupportedSource: OpenRouter leaderboard, via The New Stack2 sources— create a free account to open themView cited source - [4]
Elon Musk posted an "important note" on Wednesday about Grok Bot, the AI agent SpaceX launched in beta in August.
- [5]
"Important note regarding Grok @Bot: Going forward, @SpaceX will use the best back end model for any given task, including Claude Opus 5.5, MidJourney, Suno and other leading APIs. Whatever is most likely to give you the best outcome."
- [6]
Grok Bot is a joint product of SpaceXAI, the Musk AI lab now folded into SpaceX, and Cursor, which SpaceX agreed in June to buy for $60 billion in stock in a deal that closed in August.
- [7]
Burns wrote in August that "when models converge on price, the money moves to whoever decides which model gets the job."
- [8]
The 22 interviewees run Insight Partners portfolio companies, the interviews were shot for a video series Insight is producing, and Insight Partners owns The New Stack.
- [9]
A couple of the people interviewed sell small models for a living, and Burns said he weighs their enthusiasm accordingly.
- [10]
The most common model triage setup described was a funnel: one security company runs a rules engine over everything first, passes what survives to small models, and larger models only see what is left.
- [11]
An executive of the security company said running a petabyte of data through any model, even a small one, would cost millions of dollars.
- [12]
Another company runs requests to its agent through a cheap model to figure out what the user wants before the expensive model does anything.
- [13]
One engineering leader said his team swapped in a newer model three days before a demo for an important prospect because benchmarks said it was cheaper and just as good; the workflow broke, the team rolled back over a weekend, and his lesson was evals.
- [14]
The New Stack's Amanda Caswell reported that Anthropic made Opus 5.5 20% cheaper than Opus 5 and broke four things agents depend on.
- [15]
Anthropic's Haiku 5.5 launched Wednesday at $0.10 per million input tokens, a 90% cut for requests under 100,000 tokens.
- [16]
Small models like Haiku 5.5 typically handle high-volume jobs like classification and routing.
- [17]
Burns's advice: test before you switch, and follow the ABS rule, "Always Be Switching."
- [18]
A $0.10 per million input token price after a 90% cut implies an earlier rate of about $1.00 per million input tokens for requests under 100,000 tokens.
Sources
1 independent publisher whose own reporting we read for this story.
- thenewstack.ioElon Musk’s Grok Bot will pick Claude over Grok when it’s better. One-model loyalty is dead.
1 article · October 10, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.