Build1 publisher3 min readPublished Updated
Grok 4.6 lands in Copilot two days after launch, and the model picker becomes a procurement problem
GitHub added xAI's model to Copilot on August 14 across eight developer surfaces at usage-based pricing. The benchmark case, including xAI's own terminal scores, is mixed.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- xAI put Grok 4.6 into GitHub Copilot on August 14, extending the two-day-old reasoning model into GitHub's editors, command-line tools and cloud coding agent.
- xAI released Grok 4.6 on August 12 with an emphasis on long-running agents, coding and multi-step knowledge work.
- GitHub is rolling out Grok 4.6 to Copilot Pro, Pro+, Max, Business and Enterprise subscribers.
- Developers will be able to select Grok 4.6 in Visual Studio Code, Visual Studio, Copilot CLI, the Copilot cloud agent, the Copilot app, JetBrains IDEs, Xcode and Eclipse.
- GitHub said availability will expand gradually rather than reach every eligible account immediately.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
xAI's Grok 4.6 became selectable inside GitHub Copilot on August 14, two days after the model was released [1][2][7]. The short interval is the point: picking a model in an IDE has stopped being a vendor default and become a procurement decision, with cost, evidence quality and counterparty exposure attached to it.
GitHub is rolling the model out to Copilot Pro, Pro+, Max, Business and Enterprise subscribers, and it shows up in eight places: Visual Studio Code, Visual Studio, Copilot CLI, the Copilot cloud agent, the Copilot app, JetBrains IDEs, Xcode and Eclipse [4][5]. GitHub says availability expands gradually rather than reaching every eligible account immediately [6], which in practice means your developers will see inconsistent pickers for a while and file tickets about it. Grok 4.6 is billed at provider list pricing under Copilot's usage-based system, with GitHub's pricing table setting the applicable model and request rates [11][12]. Anyone administering a Business or Enterprise seat pool is now managing a menu whose items carry different unit costs.
The catalog itself has become the market. GitHub's supported-model documentation lists models from OpenAI, Anthropic, Google, Microsoft, Moonshot AI and xAI [8], six providers behind one interface, with GitHub holding the interface, the billing and the policy settings [9][10]. That places the question of which providers receive your prompts at the platform layer, not with the model vendor.
The evidence for switching it on is thinner than the announcement suggests. GitHub said its internal testing produced strong results on terminal-based coding tasks in VS Code and Copilot CLI, particularly in longer workflows requiring sustained reasoning and tool use, but it published no task set, error rates, sample size or direct comparisons [13][14]. xAI's own published numbers run the other way on precisely that axis: 26% on Terminal-Bench v3.0, against the 34.6% xAI listed for GPT-5.6 Sol and 34.1% for Claude Fable 5 [17], a gap of 8.6 points that leaves Grok 4.6 at roughly three quarters of the leader's score [18]. On DeepSWE v1.1, xAI reported 65.9% versus 73% for GPT-5.6 Sol and 70% for Fable 5 [16], 7.1 points off the top figure [27]. CursorBench v3.2 is the favourable one: 69.9%, ahead of GPT-5.6 Sol by 2.7 points and behind Fable 5 by 0.6 [15][28]. On the Artificial Analysis Intelligence Index, xAI reported 61, level with GPT-5.6 Sol and one point behind Fable 5 at 62 [19][20]. xAI says the competitor figures came from developer-published system cards or public leaderboards [21], and its claims about a longer supplemental training run than Grok 4.5 and more self-testing on long trajectories come from its own evaluation process [22]. Read together, that is a competitive coding model rather than a category leader, and GitHub's terminal claim sits awkwardly beside its supplier's own terminal benchmark.
There is also a counterparty fact for the vendor file. SpaceX acquired xAI on February 2, 2026, and the company now brands itself SpaceXAI while continuing to ship models as Grok [23]. Copilot's reach comes from Microsoft's $7.5 billion purchase of GitHub in 2018, which tied it to existing hosting, review and deployment workflows [25]. Enabling Grok routes developer prompts to a rocket company's subsidiary through a Microsoft-owned code host.
Three things to watch: whether GitHub publishes the terminal task set behind its assessment [14], whether per-model rates in the pricing table make Grok 4.6 materially cheaper or dearer than the alternatives once the first usage bills land [11][12], and whether enterprise administrators leave model policy on allow-all or move to an approved list.