Skip to content

Build2 publishers3 min readPublished Updated

Foundry offers Provisioned Throughput for two of the three GPT-6 models

Microsoft says evaluations should decide which GPT-6 model runs each agent. The serving matrix leaves Luna without reserved capacity. GitHub turns both new models on for admins who left the default alone.

The Engineer · Build desk

Illustration accompanying Foundry offers Provisioned Throughput for two of the three GPT-6 models

What happened

  • Microsoft made GPT-6 Sol and GPT-6 Luna generally available in Microsoft Foundry alongside GPT-6 Astra, which it recommends as the starting point for demanding reasoning and computer-use work.
  • GitHub added both models to the Copilot picker in VS Code, Visual Studio, the Copilot CLI, JetBrains, Xcode, Eclipse and its mobile apps, billed under usage-based billing on a gradual rollout.
  • Copilot Business and Enterprise admins govern access through model policy, where new models switch on automatically unless the global default is off or that model is explicitly disabled.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision A team without comparison runs ends up on whatever tier GitHub's model policy turns on for its coding agents, and the cheapest place to find that out is an eval run.
  • constraint Any high-volume step that also needs guaranteed throughput has to be promoted from Luna to Sol, so predictability comes at a higher unit cost.
  • exposure A latency-sensitive agent pinned to the EU Data Zone falls outside the region list for Priority Processing, so an EU tenant and a Global tenant running the same agent are not on the same serving tier.
  • cost Cost per task is the measure Microsoft proposes, and producing it is customer-side work: traces, token counts per tier and a completion label per task, before any migration is priced.

The setting that decides this for most teams is in Copilot's model policy. GitHub's changelog says that under default model enablement, new models are enabled automatically unless an administrator has turned off the global default or explicitly disables the model [14]. Rollout is gradual [13]. On a Business or Enterprise tenant where nobody has opened that page, Sol and Luna arrive in the picker in VS Code, Visual Studio, the Copilot CLI, JetBrains, Xcode and Eclipse on GitHub's schedule [12].

Microsoft's own guidance treats the choice as an evidence problem. "The right model for a job should be determined through evaluations: an agent handling a complex business decision and one routing routine requests have different needs," the Azure post says [4]. It recommends starting with Astra for demanding work and carrying higher-volume workloads to Sol and Luna [2], with Luna aimed at extraction, summarization, request routing and routine customer interactions [3].

Then read the deployment list, because it does not line up with the tiers. Standard covers all three models across all 28 Global regions plus the US and EU Data Zones [5]. Provisioned Throughput covers Astra and Sol, in Global regions and the US and EU Data Zones [6]. Priority Processing is listed for Sol only, in Global regions and US Data Zones [7]. Count the options per model: Sol appears in three, Astra in two, Luna in one [17]. Luna is the model Microsoft recommends for the highest call volumes [18].

The plan matrix runs the other way. Sol needs Copilot Pro+, Max, Business or Enterprise, and Luna is the one of the two that also reaches Copilot Pro [11][20]. Both bill under usage-based billing, with GitHub pointing to its pricing page for rates [15].

On money, Microsoft writes that customers should look beyond pricing per token and seek to understand cost per task, which it calls a better measure of ROI [8]. The post refers to an accompanying chart it says illustrates why enterprise customers on Foundry are switching to GPT-5.6 Sol and the latest GPT-6 offerings; the post text does not include the chart's figures [9].

The named customer in the post talks about integration and support. "Through Azure, we reliably integrate advanced models into our agentic workflows, enabling Manus to understand user intent, plan tasks, and execute complex work," said Tao Zhang, Co-Founder and Product Partner at Manus [16].

So the harness has real work to do. For each agent step you need a completion measure, token counts per tier, and a note of which serving option that step actually has available in the region it runs in. Sol is the only model in the lineup for which all three serving options are listed [17].

What to watch

  • Whether Provisioned Throughput is extended to Luna, which would make the cheapest tier viable for reserved-capacity batch work.
  • Whether Priority Processing reaches the EU Data Zone, or Astra and Luna, beyond the current Sol-only listing.
  • Whether Microsoft publishes the cost-per-task figures behind the switching chart, or per-token prices for the three models.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories