Build1 publisherNot yet confirmed elsewhere2 min readPublished
GitHub Copilot's on-device MAI Code 1.1 Flash peaks at 75.5GB of memory
Microsoft plans to run MAI Code 1.1 Flash on developers' machines in GitHub Copilot, with a limited rollout starting by end-October 2026. Its figures come from one high-end laptop and local pricing is undisclosed, so teams cannot yet size the hardware or the bill.
The Engineer · Build desk
What happened
- The first local rollout covers Copilot CLI, the Copilot app and IDE integrations such as Visual Studio Code.
- Per a GitHub Changelog entry, Microsoft had widened MAI-Code-1-Flash's availability to additional Copilot surfaces back in June 2026.
- Developers can let Copilot's Auto mode route each request or pick the model explicitly through the Windows ML provider or a local endpoint.
- Microsoft is pairing local execution with Microsoft Execution Containers, or MXC, which are meant to isolate the tools an agentic coding workflow runs.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint Memory decides which developer machines qualify, and the only published target is one high-end reference laptop, so fleet planning has to wait for a device list.
- decision Teams with data-residency rules will probably standardize on explicit local selection, since it is the mode where the developer fixes where each request is processed.
- cost Any business case that books local calls as free rests on a social post until Microsoft sets subscription treatment and usage limits.
- exposure Agentic tool calls running on developers' own machines put MXC's isolation in scope for security review before a team turns the feature on.
Microsoft's overview, as dev.to reports it, sets compute and memory separately. MAI Code 1.1 Flash has 137 billion total parameters and 6.8 billion active ones [5]. About 5% of the model is active on any given token [18]. That share sets the compute per token. Memory is set by the whole model. Microsoft's reference run peaked at about 75.5GB with the full 256,000-token context loaded [6][11]. Spread across 137 billion parameters, that peak comes to roughly 0.55 bytes each [20]. Microsoft names quantization and speculative decoding as the techniques that shrink the on-device footprint [5].
The reference machine is a Surface Laptop Ultra with NVIDIA RTX Spark [11]. Decode throughput was 923.5 tokens per second at 64,000 tokens of context and 769.8 at 128,000 [12][13]. Doubling the context cost about 17% of decode speed [19]. Those are decode rates from one configuration. For them to transfer, a developer's machine needs a comparable accelerator and enough memory for the context length the team actually uses. Smaller machines need their own measurements.
Auto mode rests on HydraFusion, Microsoft's approach for coordinating models across edge and cloud environments [7]. Cloud inference stays in the stated routing model, and dev.to describes the result as a hybrid workflow [9]. For a team that must document where source code was processed, a per-request router is one more decision to audit. Explicit selection puts that choice with the developer [17].
Pairing the local model with sandboxed tool execution is the right order of work [10]. A local model that calls tools runs them on the developer's own machine. The isolation should be in place before agentic workflows reach that machine, and Microsoft is shipping it with the local path.
According to dev.to, Microsoft's official announcement does not disclose local-inference pricing, subscription treatment or usage limits, and it does not include a minimum hardware specification or a supported-device list [4][14]. The statement that local calls carry no inference charge came from an originating social post [15].
In my context, a team with residency rules and a few high-memory workstations, I would trial explicit local selection on those machines first. Auto stays off until the price is public.
What to watch
- Microsoft's published pricing, subscription treatment and usage limits for local inference in Copilot.
- A minimum hardware specification or supported-device list, especially any configuration with less memory than the reference Surface Laptop Ultra.
- How Auto orchestration splits requests between local and cloud once the limited rollout starts by the end of October 2026.