BuildAlso reported elsewhere2 publishers3 min readPublished
Microsoft's Decision-1 turns the check before an agent acts into a scored Foundry call
Microsoft has put Decision-1, a model that scores fixed choices, in Foundry at $0.042 per million input tokens. Agent teams now have to test whether its scores, so far benchmarked only by Microsoft, are calibrated well enough to decide when a case goes to a human.
The Engineer · Build desk

What happened
- Microsoft says the model had the highest accuracy in its own comparison of 36 benchmarks and nearly 150,000 questions, with the benchmarks kept blind from training.
- Inside Microsoft, Xbox Research used it to sort more than 10,000 feedback items, and the Copilot team used it to assess AI responses.
- The launch is led by Achint Srivastava, a Microsoft vice president who co-founded the AI evaluation startup Pi Labs with David Karam.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost A check that fills the whole input window costs about a seventh of a cent, so most of the adoption cost is the labeled evaluation set each team must build from its own decisions.
- exposure Teams that point it at employment, credit, housing, healthcare or legal-rights calls keep responsibility for thresholds and human review, because Microsoft's catalog says the model should not decide those alone.
- decision Picking a Foundry-hosted scorer over Cloudflare's open-weight Clef ties escalation thresholds to Microsoft's endpoint until the undated OpenRouter access arrives, and moving later means recalibrating.
The interface is a closed question. An application sends a situation and the list of allowed answers, and Decision-1 sends back a probability for each one [2]. Input is text only, up to 32,768 tokens, and the response is JSON [4]. RuntimeWire's headline calls it a 9B model [3]. The supported formats are yes-or-no, multiple choice, ratings, rubric grading and relevance judgments, plus an explicit "cannot tell" option [5]. It does not write explanations or take open-ended questions, and it does not accept images, audio or video [6]. Srivastava's October 9 announcement lists its jobs as routing, classification, prioritization, verification and workflow control [17].
The "cannot tell" option is the best decision in the listing. A scorer that must split all its probability between "proceed" and "reject" has nowhere to put an input it does not understand. With an abstain choice available, escalation logic gets simple. RuntimeWire describes the pattern: an agent asks whether a proposed tool call meets a rubric, then proceeds, retries or escalates based on the score [7]. A team can send any case where "cannot tell" crosses a set level straight to a person.
Serial latency is the constraint behind the product. Agent checks are usually dependent, so step twelve waits on the verdict for step eleven. Microsoft's engineers give the figure in the announcement: 100 milliseconds added to each of 20 dependent decisions adds two seconds to a workflow [8]. Microsoft's median-latency figures put Decision-1 ahead of GPT-6 Sol by a factor of 35 and ahead of Quyet-1.0-Large by a factor of 4.5 [21]. A small scorer beating a general text generator on a task whose output is a few numbers is the expected direction of that result.
Input costs $0.042 per million tokens, and output is free [9]. Free output is an easy promise for a model that replies with a short list of probabilities. A check that fills the whole 32,768-token window costs about $0.0014 [19]. A thousand of those cost about $1.38 [20].
Across 36 benchmarks totaling nearly 150,000 questions, none of which the model saw in training, Microsoft says Decision-1 was the most accurate of the models it compared [22]. RuntimeWire notes that the announcement leaves out the full datasets and the latency test conditions, so outsiders cannot reproduce the comparison from it [10]. For the accuracy figure to carry over, a team's decisions have to resemble those benchmarks in their label sets, their prompt lengths and how often the right answer is actually ambiguous. For the latency multiples to carry over, the test's prompt sizes and concurrency have to look like production traffic.
Accuracy and calibration are separate properties. Accuracy counts how often the top choice is right. An escalation threshold needs a score of 0.8 on "proceed" to be right about 80 percent of the time on the team's own traffic. RuntimeWire frames the open question the same way: whether the score is dependable on customer-specific decisions, and calibrated well enough to set escalation thresholds safely [11]. A labeled sample of a team's own decisions, scored and plotted as predicted probability against observed outcome, answers that before any threshold ships.
Hosting is the other choice. Cloudflare introduced its Clef models on October 1 with open weights, deployed through Workers AI [14]. OpenAI's Decisions API is in public beta and returns typed answers to fixed questions [15]. Decision-1 starts in Foundry, and Microsoft says OpenRouter access is coming soon, without a date [16]. Thresholds tuned against one model's probabilities do not move to another model unchanged, so switching scorers means calibrating again. My context is a workflow already on Azure with a labeled eval set. There, I would trial Decision-1 first on support-ticket routing, one of the uses Microsoft names [18], because a misrouted ticket is cheap to reverse.
What to watch
- Publication of the 36 benchmark datasets or the latency test conditions, which would let outside teams reproduce Microsoft's accuracy and 35x latency claims.
- Calibration data for Decision-1, such as predicted-versus-observed curves on customer workloads, which is what escalation thresholds depend on.
- A date for Decision-1 on OpenRouter, which would loosen the tie to Foundry hosting.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence35
- Adoption12
- Hype gap+30
- Incentives70
- Confidence38
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Microsoft has put Microsoft-Decision-1, a model that scores fixed choices instead of generating text, in Microsoft Foundry.
- [2]
Given a situation and a defined set of choices, Decision-1 returns a probability for each option.
- [3]
RuntimeWire's headline describes Decision-1 as a 9B decision model.
- [4]
Microsoft's Foundry listing says the model accepts text inputs of up to 32,768 tokens and returns JSON.
- [5]
Decision-1 supports yes-or-no questions, multiple choice, ratings, rubric grading, relevance judgments and an explicit "cannot tell" option.
- [6]
Decision-1 does not generate explanations, handle open-ended questions or accept images, audio or video.
- [7]
An agent can ask whether a proposed tool call meets a rubric, then proceed, retry or escalate based on the score.
- [8]
Microsoft's engineers say in the announcement that adding 100 milliseconds to each of 20 dependent decisions adds two seconds to a workflow.
- [9]
Microsoft lists Decision-1 input tokens at $0.042 per million, with output tokens free.
- [10]
Microsoft's announcement does not provide the full datasets or latency test conditions needed to reproduce the comparison independently.
- [11]
RuntimeWire says the greater operational question is whether the score is dependable on the customer-specific decisions developers need to automate, and whether it is calibrated well enough to set escalation thresholds safely.
- [12]
The Foundry catalog warns that Decision-1 should not be the sole automated decision-maker for consequential matters involving employment, credit, housing, healthcare or legal rights; integrating applications need to set thresholds, human review and safeguards.
- [13]
Achint Srivastava, a Microsoft vice president, co-founded AI evaluation startup Pi Labs with David Karam; Accel's profile lists Microsoft as Pi Labs' acquirer. The sources do not establish that Pi Labs' technology or team built Decision-1.
- [14]
Cloudflare introduced its Clef models on October 1st, with open weights and deployment through Workers AI.
- [15]
OpenAI's Decisions API is in public beta, returning typed answers for fixed questions.
- [16]
Microsoft plans to bring Decision-1 to OpenRouter; its announcement says OpenRouter access is coming soon, without giving a date.
- [17]
In his October 9th announcement, Srivastava describes Decision-1 as a tool for routing, classification, prioritization, verification and workflow control.
- [18]
Developers can use Decision-1 scores to route a support ticket, check an agent's proposed action, or send an uncertain case for human review.
- [19]
A single Decision-1 check that fills the full 32,768-token input window costs about $0.0014 at list price.
- [20]
One thousand maximum-length Decision-1 checks cost about $1.38 at list price.
- [21]
Microsoft reports median latency 35 times faster than GPT-6 Sol and 4.5 times faster than Quyet-1.0-Large.
- [22]
Microsoft says Decision-1 had the highest accuracy in its comparison of 36 benchmarks and nearly 150,000 questions, with benchmarks kept blind from training.
- [23]
Xbox Research used Decision-1 to sort more than 10,000 feedback items, and Copilot's team used it to assess AI responses; these results are Microsoft-reported.
Sources
2 independent publishers whose own reporting we read for this story.
- runtimewire.comMicrosoft puts a 9B decision model in Foundry for agent workflows
1 article · October 9, 2026
- testingcatalog.comMicrosoft launches Decision-1 model in Foundry
1 article · October 9, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Entities
- MicrosoftFollow
- Microsoft FoundryFollow
- Microsoft-Decision-1Follow
- Achint SrivastavaFollow
- Pi LabsFollow
- David KaramFollow
- AccelFollow
- CloudflareFollow
- ClefFollow
- Workers AIFollow
- OpenAIFollow
- Decisions APIFollow
- OpenRouterFollow
- AlibabaFollow
- Qwen3.5-9BFollow