On October 9, Microsoft announced Microsoft-Decision-1, an AI model that returns a probability for each of a set of predefined options instead of generating a written answer. It is generally available on Microsoft Foundry and can also be used through OpenRouter. The model is Alibaba's Qwen3.5-9B with additional training, and it is intended for uses that repeat large volumes of decisions quickly and cheaply, such as classifying inquiries and evaluating the actions of AI agents. Microsoft claims a speed roughly 35 times that of GPT-6 Sol, but the gap with other decision-specific models is much smaller, and third-party response-time measurements differ from the official figure. Being able to evaluate options quickly does not mean business decisions can simply be handed over to it.
Built on Qwen, but specialized for decisions rather than text generation
Input to Decision-1 consists of state, which specifies the information needed for the decision, and questions, which asks about that information. According to Microsoft's developer documentation, both text and JSON can be used as input data, and multiple questions can be sent at once about the same information. The questions themselves can be written in natural language, but the model returns numerical decision results rather than explanatory text.
For example, to route a customer inquiry to the right department, you would prepare options such as "billing," "technical," "account," and "needs review," and specify the scope of each. The model returns the option it judges most appropriate, along with the probability that each option applies. Actually forwarding or replying to the inquiry is done by the application that receives the result. This is an illustration of how the mechanism works, not a deployment case published by Microsoft.
The API supports three types: noul (true/false judgment), which returns the probability that a condition holds; choice (selection), which picks one from multiple candidates; and score (evaluation), which grades against graduated criteria. It can also evaluate AI-generated answers or proposed agent actions against preset criteria.
This allows a division of labor in which a generative AI drafts an answer and Decision-1 evaluates it.
The base is Qwen3.5-9B, a model with 9 billion parameters. Microsoft says it performed additional training using public and synthetic data, tuning the model to evaluate each option in a single inference pass. It is an existing language model optimized for decision processing, not a fundamentally new AI technology distinct from conventional LLMs.
Nor can all of the original Qwen3.5-9B's capabilities be used. Qwen3.5-9B can handle images, but the Decision-1 model card states explicitly that it is a text-only model.
The input limit is 32,768 tokens, and free-form conversation, translation, and summarization are not intended uses. Details such as how the model evaluates options internally and what methods were used in the additional training have not been disclosed.
What does "35x faster than GPT-6 Sol" actually compare?
The performance comparison in Microsoft's announcement needs to be read as three separate metrics: accuracy, response time, and probability calibration.
Accuracy was evaluated on 36 benchmarks that Microsoft says were not used in training, totaling 147,137 questions. The following table extracts the main models from the comparison table Microsoft published.
| Model | Average accuracy | Median response time (p50) | Probability calibration score |
|---|---|---|---|
| Microsoft-Decision-1 | 83.5% | 85 ms | 92.2 |
| Jev 1.13.0 | 82.3% | 240 ms | 93.7 |
| Quyet-1.0-Large | 81.9% | 380 ms | 93.1 |
| GPT-6 Luna Decisions | 79.4% | 300 ms | 89.9 |
| GPT-6 Sol | Not listed | 3,010 ms | Not listed |
Accuracy and the probability calibration score are based on Microsoft's evaluation across the 36 benchmarks. The calibration score indicates how closely the probabilities the model reports match its actual accuracy, with 100 meaning a perfect match. It is not itself a measure of accuracy.
As for response time, a footnote says Microsoft used the adjusted median from JevBench v1.6.1, checked on October 7, 2026. However, Decision-1's 85 ms is a value measured within the same region on Microsoft Foundry. The three types of metrics in the table were therefore not all measured in tests under identical conditions.
Calculated from the published response times, Decision-1 is about 35.4 times as fast as GPT-6 Sol and about 3.5 times as fast as GPT-6 Luna Decisions, a decision-specific model.
The calculations are 3,010 ÷ 85 ≈ 35.4 and 300 ÷ 85 ≈ 3.5.
In other words, the "roughly 35x" that Microsoft emphasizes is a comparison with the general-purpose GPT-6 Sol, not with a model that also specializes in decision processing. It is also a simple comparison of published values obtained under different measurement conditions, not the result of measuring all models simultaneously in the same environment.
Nor does it mean an entire application becomes 35 times faster. Time spent on non-decision processing, such as search and running external tools, must be considered separately.
Decision-1 ranked first in accuracy but third in probability calibration, behind Jev and Quyet.
The ability to choose the correct option and the ability to express, with an appropriate probability, how likely one's own judgment is to be correct are different things. Since GPT-6 Sol's accuracy is not listed in the comparison table, it cannot be concluded that Decision-1 also beats the general-purpose model in decision accuracy on the strength of its speed.
Third-party measurement gives 459 ms versus the official 85 ms
JevBench from Benchmark Heaven, which independently evaluates decision-specific models, also added results for Decision-1 on October 10.
This evaluation used 1,200 private questions and 300 public questions sent directly to Azure Foundry. In the 124 requests used to measure response speed, the median response time was 459 ms. The 85 ms published by Microsoft therefore cannot be taken as-is as the response time in actual usage environments.
In the API ranking as of the same day, Decision-1 placed 6th among 27 models with an overall score of 69.11.
However, the figure 69.11 does not mean 69.11% accuracy. It is JevBench's own composite metric, which weights decision ability and probability calibration together with processing speed and price. The questions used and the aggregation method also differ from Microsoft's 36-benchmark evaluation.
JevBench also counts 21 failures caused by inputs exceeding the context limit, rather than excluding them from the evaluation.
Furthermore, for other models that were already listed, it keeps the measurement dates and question samples from that time. Because all models were not re-measured using this question set, it is difficult to compare model performance uniformly from the ranking alone.
That said, the gap between 85 ms and 459 ms alone does not allow one to conclude that Microsoft's announcement is wrong or that the model's performance has declined. The measurement dates, network paths, and question contents differ.
The 85 ms is an official value Microsoft measured within a single region, while the 459 ms is a value a third party measured through the API. Anyone deploying the model should verify response times with their own input data and connection environment.
$42 per million decisions: pricing suited to repetitive workloads
Decision-1 costs 0.042 US dollars per million input tokens, with no charge for output. The price listed on OpenRouter is the same.
For example, suppose each request has 1,000 input tokens. One million requests would total 1 billion tokens, and the input cost can be calculated as follows.
$0.042 × 1,000 = $42
This is a simple calculation based on assumed input length and request count; it does not include operating costs of surrounding systems or human review work.
In an internal Xbox Research test that Microsoft described, more than 10,000 free-text responses, such as survey answers and game reviews, were classified into themes set by researchers.
According to the company, Decision-1 maintained quality comparable to GPT-6 Sol while achieving more than 14 times the speed at one-two-hundredth of the cost. The result has not been reproduced by a third party, but for processing that repeats predetermined classifications at scale, skipping the step of having a generative AI write text each time can be expected to improve both speed and cost.
However, the actual benefit depends on how much of the overall workflow's time and cost is taken up by decision processing.
For example, in a system that performs lengthy searches and answer generation after routing an inquiry, speeding up only the routing leaves the time for the remaining steps unchanged. In a system that repeats similar decisions many times, on the other hand, even a small per-call time saving can add up to a large overall effect.
Being built on Qwen also does not mean Decision-1's own model weights are released or that it can be run in a local environment. What Microsoft is offering is an API used via Foundry or OpenRouter.
By contrast, Strands Decider 2B releases not only its model weights but also its training data and scripts, and is intended to run on local CPUs and GPUs.
Even among decision-specific models, there is a difference between using an API provided in the cloud and running and tuning the model in your own environment.
On Foundry, selecting DataZoneStandard in a supported region means inference is processed within the specified data zone. With GlobalStandard, on the other hand, compute resources in supported Azure regions are used.
Microsoft explains that choosing a nearby data zone does not necessarily mean lower latency. Before deployment, it is worth confirming not only where the company that developed the model is based, but also where the data you submit is actually processed.
Even with probabilities, how far to automate is the user's call
If you collect many decisions for which the model returned "90%" and about 90% of them turn out to be correct, the predicted probability and actual accuracy can be said to match well. Improving that match is what probability calibration does.
Even if accuracy is high, a model that assigns high probabilities to wrong decisions makes it hard to use those numbers as a basis for splitting work between automated processing and human review.
Microsoft also measured how stable the decisions are against changes in input. When instructions, options, and other elements were altered in eight different ways, the proportion of decisions that changed averaged 1.3%, it reports.
However, this does not guarantee stable handling of any rephrasing in actual usage environments. The developer documentation itself notes that scores may change depending on how a question is worded or the order of options, and asks users to verify with data representative of their actual use.
Although Decision-1 does not generate text, it can still choose the wrong option.
For example, if an inquiry that should have been checked as a refund matter is routed to technical support, the business decision is wrong even if the output JSON is formally valid. And if the correct answer is not among the options to begin with, no matter how clear the returned numbers are, it cannot be said that a correct decision was made.
For that reason, it is important to provide options for withholding judgment, such as "unclear" or "needs review."
The threshold for what probability is sufficient to proceed with automated processing must also be set on the application side.
The appropriate threshold differs between work where a wrong decision causes large losses and work where some misclassifications can be corrected later. You need to attach correct labels to the Japanese-language inquiries you actually handle and to hard-to-classify cases, investigate what kinds of wrong decisions occur, and then set the conditions under which a human or another model takes over the check.
The score type for grading also requires care.
For example, in a four-level evaluation, each level is assigned a number from 0 to 3, and the score the model returns is a weighted average calculated using the probability of each level.
Even if a value falling between levels is returned, that does not mean it is an objective quality score valid for every purpose. Microsoft also recommends using it for ranking and threshold-based judgments rather than treating it as an absolute grade.
Furthermore, Decision-1 has no function for explaining the reasons for its decisions in text. Microsoft asks that the model not be used as the sole means of making decisions with serious consequences for individuals, such as in healthcare and employment.
OpenRouter also explains that it will continue to update the model weights while maintaining the API format. Thresholds set at deployment therefore need to be checked periodically to confirm they still work the same way after model updates.
Microsoft has also revealed plans to switch the underlying model to one of its own MAI models or an OpenAI model.
Even if the base model changes, what matters is that it keeps returning stable decisions on your own data and can properly switch between automated processing and human review according to probability. If such a mechanism can be established, developers will be able to use models suited to text generation and models that repeat large volumes of decisions at high speed, depending on the purpose of the work.
