Build1 distinct publisher3 min readUpdated
GitHub added xAI's model to Copilot on August 14 across eight developer surfaces at usage-based pricing. The benchmark case, including xAI's own terminal scores, is mixed.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
GitHub added xAI's model to Copilot on August 14 across eight developer surfaces at usage-based pricing. The benchmark case, including xAI's own terminal scores, is mixed.
xAI's Grok 4.6 became selectable inside GitHub Copilot on August 14, two days after the model was released [1][2][7]. The short interval is the point: picking a model in an IDE has stopped being a vendor default and become a procurement decision, with cost, evidence quality and counterparty exposure attached to it.
GitHub is rolling the model out to Copilot Pro, Pro+, Max, Business and Enterprise subscribers, and it shows up in eight places: Visual Studio Code, Visual Studio, Copilot CLI, the Copilot cloud agent, the Copilot app, JetBrains IDEs, Xcode and Eclipse [4][5]. GitHub says availability expands gradually rather than reaching every eligible account immediately [6], which in practice means your developers will see inconsistent pickers for a while and file tickets about it. Grok 4.6 is billed at provider list pricing under Copilot's usage-based system, with GitHub's pricing table setting the applicable model and request rates [11][12]. Anyone administering a Business or Enterprise seat pool is now managing a menu whose items carry different unit costs.
The catalog itself has become the market. GitHub's supported-model documentation lists models from OpenAI, Anthropic, Google, Microsoft, Moonshot AI and xAI [8], six providers behind one interface, with GitHub holding the interface, the billing and the policy settings [9][10]. That places the question of which providers receive your prompts at the platform layer, not with the model vendor.
The evidence for switching it on is thinner than the announcement suggests. GitHub said its internal testing produced strong results on terminal-based coding tasks in VS Code and Copilot CLI, particularly in longer workflows requiring sustained reasoning and tool use, but it published no task set, error rates, sample size or direct comparisons [13][14]. xAI's own published numbers run the other way on precisely that axis: 26% on Terminal-Bench v3.0, against the 34.6% xAI listed for GPT-5.6 Sol and 34.1% for Claude Fable 5 [17], a gap of 8.6 points that leaves Grok 4.6 at roughly three quarters of the leader's score [18]. On DeepSWE v1.1, xAI reported 65.9% versus 73% for GPT-5.6 Sol and 70% for Fable 5 [16], 7.1 points off the top figure [27]. CursorBench v3.2 is the favourable one: 69.9%, ahead of GPT-5.6 Sol by 2.7 points and behind Fable 5 by 0.6 [15][28]. On the Artificial Analysis Intelligence Index, xAI reported 61, level with GPT-5.6 Sol and one point behind Fable 5 at 62 [19][20]. xAI says the competitor figures came from developer-published system cards or public leaderboards [21], and its claims about a longer supplemental training run than Grok 4.5 and more self-testing on long trajectories come from its own evaluation process [22]. Read together, that is a competitive coding model rather than a category leader, and GitHub's terminal claim sits awkwardly beside its supplier's own terminal benchmark.
There is also a counterparty fact for the vendor file. SpaceX acquired xAI on February 2, 2026, and the company now brands itself SpaceXAI while continuing to ship models as Grok [23]. Copilot's reach comes from Microsoft's $7.5 billion purchase of GitHub in 2018, which tied it to existing hosting, review and deployment workflows [25]. Enabling Grok routes developer prompts to a rocket company's subsidiary through a Microsoft-owned code host.
Three things to watch: whether GitHub publishes the terminal task set behind its assessment [14], whether per-model rates in the pricing table make Grok 4.6 materially cheaper or dearer than the alternatives once the first usage bills land [11][12], and whether enterprise administrators leave model policy on allow-all or move to an approved list.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
xAI put Grok 4.6 into GitHub Copilot on August 14, extending the two-day-old reasoning model into GitHub's editors, command-line tools and cloud coding agent.
xAI reported a 26% score for Grok 4.6 on Terminal-Bench v3.0, below the 34.6% and 34.1% figures it published for GPT-5.6 Sol and Claude Fable 5.
xAI released Grok 4.6 on August 12 with an emphasis on long-running agents, coding and multi-step knowledge work.
GitHub is rolling out Grok 4.6 to Copilot Pro, Pro+, Max, Business and Enterprise subscribers.
Developers will be able to select Grok 4.6 in Visual Studio Code, Visual Studio, Copilot CLI, the Copilot cloud agent, the Copilot app, JetBrains IDEs, Xcode and Eclipse.
GitHub said availability will expand gradually rather than reach every eligible account immediately.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Verifiable distribution facts, provider-graded capability
Distribution details — date, tiers, surfaces, billing model, default-off enterprise policy — are specific and checkable against GitHub's own documentation. The capability case is much weaker: every comparative score is published by xAI, competitor figures are lifted from system cards or leaderboards rather than rerun, and GitHub's supporting internal testing arrives with no task set, error rates, sample size or comparisons. One publisher carries the whole cluster, so nothing is independently corroborated.
Broad availability, zero demonstrated usage
Availability is genuinely wide — five paid tiers and eight developer surfaces — and the pricing table already lists Grok 4.6 at Grok 4.5 rates, so the plumbing is live. But adoption in the meaningful sense is unmeasured: the rollout is gradual rather than universal, enterprise access is off by default until an administrator flips it, and neither GitHub nor xAI has published selection counts, session volumes or agent-run lengths. Availability is being scored, not uptake.
Terminal claim outruns the terminal score
The mismatch is specific rather than diffuse. GitHub frames the addition around strong terminal-based coding in VS Code and Copilot CLI, yet xAI's own Terminal-Bench v3.0 figure for Grok 4.6 is 26% against 34.6% and 34.1% for the two rivals it lists — the weakest of the four published comparisons. Grok 4.6 leads on CursorBench v3.2 and ties GPT-5.6 Sol on the Artificial Analysis index, so this is a moderately overstated positioning, not an unsupported launch. The source itself flags the tension, which limits the gap.
Every capability number is vendor-supplied
Both parties benefit from the framing. xAI publishes the comparison tables and the training-run narrative, and gains distribution into GitHub's installed base without displacing anyone's toolchain; capital from a reported $20B Series E and the SpaceX acquisition depends on exactly this kind of placement. GitHub and Microsoft benefit from another provider competing on price and capability inside a picker whose interface, billing and policy GitHub controls, and supply the only supporting capability evidence — unpublished internal testing. No independent evaluator appears anywhere in the cluster.
Facts firm, significance unproven
Confidence in what happened is reasonably high: the availability, tiers, surfaces, billing mechanism and policy default are stated precisely enough to verify against GitHub's documentation. Confidence in what it means is low — one publisher, provider-graded benchmarks, no independent evaluation, and no usage data on either side. The mixed benchmark picture also makes the durable outcome genuinely uncertain rather than merely unreported.
build
Three frontier launches in a day, all pitched on price. Open weights set the ceiling.4 distinct publishers
science
GLM-5.3 says the quiet part: the base model did not change, the post-training did1 distinct publisher
product
Wu says Cognition is not for sale. The more useful fact is who bought Cursor last week.1 distinct publisher
leadership
The AI bill nobody reconciles: cost per finished task, not per million tokens1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 14, 2026