Skip to content

BuildAlso reported elsewhere2 publishers3 min readPublished

Microsoft's Decision-1 turns the check before an agent acts into a scored Foundry call

Microsoft has put Decision-1, a model that scores fixed choices, in Foundry at $0.042 per million input tokens. Agent teams now have to test whether its scores, so far benchmarked only by Microsoft, are calibrated well enough to decide when a case goes to a human.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Microsoft's Decision-1 turns the check before an agent acts into a scored Foundry call
Generated illustration

What happened

  • Microsoft says the model had the highest accuracy in its own comparison of 36 benchmarks and nearly 150,000 questions, with the benchmarks kept blind from training.
  • Inside Microsoft, Xbox Research used it to sort more than 10,000 feedback items, and the Copilot team used it to assess AI responses.
  • The launch is led by Achint Srivastava, a Microsoft vice president who co-founded the AI evaluation startup Pi Labs with David Karam.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost A check that fills the whole input window costs about a seventh of a cent, so most of the adoption cost is the labeled evaluation set each team must build from its own decisions.
  • exposure Teams that point it at employment, credit, housing, healthcare or legal-rights calls keep responsibility for thresholds and human review, because Microsoft's catalog says the model should not decide those alone.
  • decision Picking a Foundry-hosted scorer over Cloudflare's open-weight Clef ties escalation thresholds to Microsoft's endpoint until the undated OpenRouter access arrives, and moving later means recalibrating.

The interface is a closed question. An application sends a situation and the list of allowed answers, and Decision-1 sends back a probability for each one [2]. Input is text only, up to 32,768 tokens, and the response is JSON [4]. RuntimeWire's headline calls it a 9B model [3]. The supported formats are yes-or-no, multiple choice, ratings, rubric grading and relevance judgments, plus an explicit "cannot tell" option [5]. It does not write explanations or take open-ended questions, and it does not accept images, audio or video [6]. Srivastava's October 9 announcement lists its jobs as routing, classification, prioritization, verification and workflow control [17].

The "cannot tell" option is the best decision in the listing. A scorer that must split all its probability between "proceed" and "reject" has nowhere to put an input it does not understand. With an abstain choice available, escalation logic gets simple. RuntimeWire describes the pattern: an agent asks whether a proposed tool call meets a rubric, then proceeds, retries or escalates based on the score [7]. A team can send any case where "cannot tell" crosses a set level straight to a person.

Serial latency is the constraint behind the product. Agent checks are usually dependent, so step twelve waits on the verdict for step eleven. Microsoft's engineers give the figure in the announcement: 100 milliseconds added to each of 20 dependent decisions adds two seconds to a workflow [8]. Microsoft's median-latency figures put Decision-1 ahead of GPT-6 Sol by a factor of 35 and ahead of Quyet-1.0-Large by a factor of 4.5 [21]. A small scorer beating a general text generator on a task whose output is a few numbers is the expected direction of that result.

Input costs $0.042 per million tokens, and output is free [9]. Free output is an easy promise for a model that replies with a short list of probabilities. A check that fills the whole 32,768-token window costs about $0.0014 [19]. A thousand of those cost about $1.38 [20].

Across 36 benchmarks totaling nearly 150,000 questions, none of which the model saw in training, Microsoft says Decision-1 was the most accurate of the models it compared [22]. RuntimeWire notes that the announcement leaves out the full datasets and the latency test conditions, so outsiders cannot reproduce the comparison from it [10]. For the accuracy figure to carry over, a team's decisions have to resemble those benchmarks in their label sets, their prompt lengths and how often the right answer is actually ambiguous. For the latency multiples to carry over, the test's prompt sizes and concurrency have to look like production traffic.

Accuracy and calibration are separate properties. Accuracy counts how often the top choice is right. An escalation threshold needs a score of 0.8 on "proceed" to be right about 80 percent of the time on the team's own traffic. RuntimeWire frames the open question the same way: whether the score is dependable on customer-specific decisions, and calibrated well enough to set escalation thresholds safely [11]. A labeled sample of a team's own decisions, scored and plotted as predicted probability against observed outcome, answers that before any threshold ships.

Hosting is the other choice. Cloudflare introduced its Clef models on October 1 with open weights, deployed through Workers AI [14]. OpenAI's Decisions API is in public beta and returns typed answers to fixed questions [15]. Decision-1 starts in Foundry, and Microsoft says OpenRouter access is coming soon, without a date [16]. Thresholds tuned against one model's probabilities do not move to another model unchanged, so switching scorers means calibrating again. My context is a workflow already on Azure with a labeled eval set. There, I would trial Decision-1 first on support-ticket routing, one of the uses Microsoft names [18], because a misrouted ticket is cheap to reverse.

What to watch

  • Publication of the 36 benchmark datasets or the latency test conditions, which would let outside teams reproduce Microsoft's accuracy and 35x latency claims.
  • Calibration data for Decision-1, such as predicted-versus-observed curves on customer workloads, which is what escalation thresholds depend on.
  • A date for Decision-1 on OpenRouter, which would loosen the tie to Foundry hosting.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence35
Adoption12
Hype gap+30
Incentives70
Confidence38
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Microsoft has put Microsoft-Decision-1, a model that scores fixed choices instead of generating text, in Microsoft Foundry.

    ReportedSupportedSource: RuntimeWireView cited source
  2. [2]

    Given a situation and a defined set of choices, Decision-1 returns a probability for each option.

    ReportedSupportedSource: RuntimeWire, describing Srivastava's announcementView cited source
  3. [3]

    RuntimeWire's headline describes Decision-1 as a 9B decision model.

    ReportedSupportedSource: RuntimeWire headlineView cited source

Sources

2 independent publishers whose own reporting we read for this story.

  1. runtimewire.com

    1 article · October 9, 2026

    Microsoft puts a 9B decision model in Foundry for agent workflows
  2. testingcatalog.com

    1 article · October 9, 2026

    Microsoft launches Decision-1 model in Foundry

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories