Skip to content

Build1 publisher3 min readPublished

Microsoft's MAI-Thinking-1 lands in Foundry, and .NET teams get a reasoning model without Python

MAI-Thinking-1 is in public preview in Microsoft Foundry at $2 per million input tokens. For C# shops the consequence is an interface swap, not a new runtime to operate.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • MAI-Thinking-1, described as Microsoft's first reasoning model, went into public preview in Microsoft Foundry, built from the ground up for multi-step reasoning including planning, reconsidering, calling tools, and stitching together long chains of context.
  • MAI-Thinking-1 slots into the same IChatClient interface already used for GPT-4o or any other Foundry model, so swapping models is a config change rather than a rewrite.
  • The described setup is: dotnet new console, add packages Azure.AI.OpenAI, Microsoft.Extensions.AI and Azure.Identity, then dotnet user-secrets init and set AZURE_AI_ENDPOINT to the Foundry resource endpoint.
  • The sample constructs the client as new AzureOpenAIClient(endpoint, new DefaultAzureCredential()).GetChatClient(deploymentName).AsIChatClient(), assigned to an IChatClient.
  • MAI-Thinking-1 is deployed from the Foundry Model Catalog to a Foundry project the same way as any other model, after which the endpoint and deployment name are used by the client.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Microsoft has put MAI-Thinking-1, described as its first reasoning model, into public preview in Microsoft Foundry [1]. The consequence for anyone shipping C# is narrower and more useful than the launch framing: the model is reachable through the same `IChatClient` interface teams already use for GPT-4o and other Foundry models, so adopting it is a configuration change rather than a rewrite [2].

The walkthrough comes from a dev.to post by Taswar Bhatti, and the setup it describes is unremarkable in the good sense. A console project, three packages (`Azure.AI.OpenAI`, `Microsoft.Extensions.AI`, `Azure.Identity`), and the endpoint held in `dotnet user-secrets` [3]. The client is an `AzureOpenAIClient` built with `DefaultAzureCredential`, then `GetChatClient(deploymentName).AsIChatClient()` [4]. The model is deployed from the Foundry Model Catalog into a Foundry project the same way as any other model, and the sample defaults the deployment name to `mai-thinking-1` [5][6]. Nothing in that chain requires a Python service, a notebook, or a second deployment target, which is the actual removal of work here.

On the model itself, the post states three things worth separating from the tooling. It uses a mixture-of-experts architecture that activates only the parts of the model a request needs, rather than the whole thing on every call [7]. It was trained from scratch on clean data rather than distilled from a third-party model, which the author frames as a provenance answer for procurement [8]. And it is described as competitive on SWE-Bench Pro at a lower price point than other models in its weight class [9]. That last claim arrives without a published score, and the post also gives no context window, latency, or rate limit figures [10], so treat it as a reason to run your own evaluation rather than a result.

Pricing is $2 per million input tokens and $8 per million output tokens through the Foundry Model Catalog [11]. Output therefore costs four times input [12], which matters more for a reasoning model than a chat model: the class of work being sold here is long chains of intermediate steps. The post does not say whether reasoning tokens are billed as output [10], and that single detail will decide whether an always-on reasoning agent is a rounding error or a budget item.

The workloads the author points at are the ones where step-by-step behaviour is load-bearing rather than decorative: agents that call CRM, ERP, or ticketing systems and must reason across the results; long-document analysis over contracts, filings, and transcripts; and decision-support features such as root-cause analysis [13]. The demonstration is a contract risk triage service whose system prompt instructs the model to identify risky clauses, reason about why each is risky, rank by severity, and emit a structured summary without skipping the reasoning steps [14].

That prompt is also the tell. In the sample, reasoning depth is controlled by telling the model in English to show its steps [14], not by a dedicated parameter in the .NET surface. The abstraction that makes model swapping cheap is the same abstraction that hides model-specific controls, so watch what `Microsoft.Extensions.AI` exposes as reasoning knobs firm up, whether token accounting distinguishes thinking from answering, and what happens to per-call latency and cost when a mixture-of-experts model is asked to plan and reconsider inside a synchronous request path.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories