Microsoft-Decision-1 is a small model from Microsoft built to pick one answer from a fixed list, fast and cheaply. CEO Satya Nadella announced it on October 9, 2026, saying Microsoft is testing it across the company for everything from incident response and quality control to scientific discovery. According to reporting that read Microsoft's announcement, it costs $0.042 per million input tokens with free output, is available in Microsoft Foundry now, and is claimed to be the fastest and most accurate in Microsoft's own comparison.
That puts Microsoft in a category that has filled up quickly this month. We have covered what decision models are, a ranking of the main ones, and launches from Perplexity, OpenAI and Liquid AI. This post explains what Microsoft's entry does, what is only a claim, and when it is worth testing. Sources: Nadella's announcement on X, the detailed read in FourWeekMBA's write-up of Microsoft's post, and Seeking Alpha's news summary. We could not open Microsoft's own post directly, so the figures below are as relayed by those outlets.
TL;DR: questions people are asking
| Question | Answer |
|---|---|
| What does it do? | Takes a question plus fixed options and returns a calibrated probability for each, in one pass. |
| What is it built on? | Qwen3.5-9B, post-trained by Microsoft. Rebasing on MAI and OpenAI models is planned. |
| Where can I get it? | Microsoft Foundry now; OpenRouter "coming soon," no date. |
| Price? | $0.042 per million input tokens; output tokens free. |
| Is it open weights? | Not stated in the coverage we found. Treat it as a hosted model. |
| Speed claim? | 2.5x faster than the runner-up (H2O-Lightning-4B v1.1) and 35x faster than GPT-6 Sol, per Microsoft. |
| Accuracy claim? | Highest accuracy in a 36-benchmark comparison, nearly 150,000 questions, kept blind from training, per Microsoft. |
| Independently tested? | No third-party results found yet. |
What does a decision model actually do?
A normal language model produces text one token at a time, and you parse the result. If you only need "which of these five queues does this support ticket belong in," that is wasteful: the model reasons in prose, spends output tokens, and may format its answer unpredictably.
A decision model skips generation. It reads the input and the options and scores each option directly. Microsoft says Decision-1 supports yes/no, multiple-choice and rating formats, plus rubric-based grading of AI responses and agent actions, through a structured API call. Because the output is a probability per option rather than free text, you can set thresholds: auto-route above 90 percent, send to a human below that.
"Calibrated" is the important word. A calibrated model's 80 percent really does mean it is right about 80 percent of the time. That is what makes the scores usable for automation, as opposed to a raw confidence number that is always overconfident. We cover the concept in more depth in our decision models guide and the supporting runtime work in llama.cpp's decision-model support.
Two options on matching plinths joined by a dotted line, showing how a decision model scores fixed choices against each other
What did Microsoft say about speed and quality?
Per the reporting on Microsoft's post, the company's headline claims are:
- Accuracy: highest in a 36-benchmark comparison covering nearly 150,000 questions kept blind from training.
- Speed: the fastest measured, 2.5 times quicker than the runner-up, H2O-Lightning-4B v1.1, and 35 times quicker than GPT-6 Sol (P50 latency).
- Robustness: when the same request was perturbed in eight ways, the model changed its decision on 1.3 percent of perturbations on average, with zero flips when option descriptions were paraphrased or when options were reversed or shuffled.
- Safety testing: 5,250 safety requests across 11 benchmarks.
- Comparison set: Quyet-1.0-Large, Surogate Rune 26B-A4B, GPT-6 Luna Decisions, deck-31B, H2O-Lightning-4B and Strands-Decider 2B.
Microsoft also gave internal results. Xbox Research used the model to sort more than 10,000 pieces of feedback and reviews into fixed themes and found it competitive with GPT-6 Sol on quality while running over 14 times faster and 200 times less expensive. The Copilot team found it competitive with GPT-5.6 Luna and 100 times faster for measuring response quality. In Microsoft Discovery it was 46 times more consistent than an LLM-based score at three times the speed.
All of these are Microsoft's own measurements. "Competitive with" is not "equal to," and the comparison set was chosen by the vendor. Several named competitors, including the open Strands-Decider 2B, are small enough to run on your own hardware, which changes the cost math in ways a per-token price does not capture.
A gauge with a lightning bolt at the needle tip, representing the claimed speed advantage of Microsoft-Decision-1
Update (October 10, 2026): what Microsoft's own post adds
We have now read the full text of Microsoft's post, Introducing Microsoft-Decision-1, by Achint Srivastava, VP of Software Engineering in Microsoft's Office of the CTO, dated October 9, 2026. It corrects one number in our earlier summary and adds detail.
- Correction. The runner-up on speed is H2O-Lightning-4B v1.1, and Decision-1 is 2.5 times quicker than it, not 4.5 times quicker than Quyet-1.0-Large as an earlier version of this post said. The 35 times figure against GPT-6 Sol stands.
- Editor's note. Microsoft says it updated the post after publication "to add benchmarks for Jev on accuracy and calibration." The accuracy comparison spans 36 benchmarks and nearly 150,000 questions that were kept blind from training. Microsoft also took several top public models from the open JevBench leaderboard and tested them across 36 additional public and private benchmarks.
- How it was built. Microsoft post-trained Qwen3.5-9B for fast, single-pass decision scoring and says it will "soon rebase it on other models, including Microsoft AI (MAI) and OpenAI." Given a fixed set of options, it returns a calibrated probability for each, and supports yes/no, multiple-choice and rating options plus rubric-based grading of AI responses and agent actions.
- Why latency compounds. The post's example: adding 100 milliseconds to each of 20 sequential decisions adds two seconds to a workflow.
- Five stated design goals. Speed, quality that generalizes beyond training tasks, robustness to paraphrase and reordering, calibrated probabilities (a 90 percent prediction should be right about nine times in 10), and safety.
- Internal use cases. Xbox Research labeling more than 10,000 feedback items (competitive with GPT-6 Sol, over 14 times faster and 200 times cheaper), Copilot quality control (competitive with GPT-5.6 Luna and 100 times faster), on-call incident knowledge retrieval (better and faster than an LLM), and Microsoft Discovery adaptive replanning (46 times more consistent than an LLM-based score at three times the speed, with nearly four times the speed on adaptive replanning).
- Availability. Microsoft Foundry and OpenRouter. Pricing is $0.042 per million input tokens, with output tokens free.
Microsoft lists a long set of suggested uses, including agent controls (continue, stop, retry or hand off), model routing, AI judging, data validation, search relevance, safety screening, computer and UI use, and robotics. As before, every performance figure is Microsoft's own, and the comparison set was chosen by the vendor.
What does it cost, and how does that compare?
At $0.042 per million input tokens with free output, a workload that classifies a million short tickets of about 300 tokens each would be roughly 300 million input tokens, or about $12.60 at list price, by our arithmetic. That is a back-of-envelope figure using Microsoft's published rate; real inputs include your options and rubric, which count as input too.
The cheaper comparison is not a frontier LLM but the alternatives in this category. The pricing of rivals is covered in our posts on Perplexity's decider and Liquid AI's open D1 models. If you can self-host a small open model, your marginal cost is compute you already own. Decision-1's pitch is that you get a managed, accuracy-leading option inside Azure's billing and compliance perimeter.
Where it fits in an agent stack
Microsoft lists agent controls among the uses: evaluating a proposed next step and deciding whether to continue, stop, retry or hand off. That is the most interesting use. Agent frameworks often burn a large model call on every micro-decision. Replacing those with a cheap scoring call is how teams cut cost and latency without changing the main model.
It ties to Microsoft's broader agent work, such as the ThinkingBox agent reliability benchmark, and to the idea of model-based grading of agent actions, which is a safety control as much as a cost control. A verifier that can say "this action looks wrong" with a calibrated probability is a building block for the guardrails we discuss in our agent safety coverage.
A green bead balanced on a curved rail between a feather and a target disc, illustrating the latency versus accuracy trade-off
Why is Microsoft shipping a decision model now?
Three pressures meet here. The first is cost: Microsoft's own announcement, as quoted by FourWeekMBA, says that with agentic AI now a reality, cost plays a major role in how people decide to use AI. Agents make dozens of small calls per task, and most of those calls are choices, not essays.
The second is competition. Within a few weeks, Perplexity, OpenAI, Liquid AI, Cloudflare and independent developers all released models or APIs that return typed decisions with probabilities. Our Cloudflare Clef Omni coverage and the Security-One 27B prompt-injection triage model show the pattern spreading from general routing into specialized domains. A hyperscaler that sells models through Foundry cannot leave the category to others.
The third is distribution. Microsoft says the model is coming to OpenRouter, which would put it in front of developers who never touch Azure. It also plans to rebase the model on MAI and OpenAI models, which suggests the recipe, not the Qwen base, is the product. That is a notable choice for a company that has been building its own models, as in our coverage of MAI-Code-1.1 Flash.
For builders, the takeaway is that the decision layer is becoming a commodity with a price floor near a few cents per million tokens. The differentiators will be calibration quality on your data, latency under load, and how well the vendor handles your compliance needs, not the headline multiples.
How to try it this week
- Pick one bounded task you currently send to a large model: ticket routing, content labeling, rubric grading.
- Build a labeled set of 200 to 500 examples from your own data. Vendor benchmarks do not predict your domain.
- Run Decision-1 and your current model side by side and compare accuracy, latency, and cost per thousand calls.
- Check calibration: bucket predictions by confidence and see whether 90 percent predictions are right about 90 percent of the time.
- Set a threshold with a human fallback and measure how many cases fall into the human bucket.
- Re-test with perturbed inputs (typos, reordered options), since robustness was a Microsoft claim worth checking yourself.
What is not established
- Microsoft's post is the only source for the benchmark results; we found no third-party evaluation.
- The OpenRouter launch has no date.
- Whether weights are downloadable is not stated in what we read.
- Quyet-1.0-Large and Surogate Rune are named as comparators, but we did not independently verify them.
- "35x faster than GPT-6 Sol" compares a single-pass scoring model with a general reasoning model on decision tasks; it is not a comparison on open-ended work.
Details are accurate as of October 9, 2026 and may change when Microsoft's documentation and independent tests appear.
