Microsoft Decision-1: Make decisions only; do not write answers.

Microsoft has launched Microsoft-Decision-1, post-trained from Qwen3.5-9B, with a focus on classification, routing, and workflow control. Rather than generating long answers, it is designed to make frequent decisions for AI systems with low latency and low cost.
Microsoft Adds a “Decision Chip” to AI Workflows
Microsoft released its dedicated decision model, Microsoft-Decision-1, on October 9 local time, and disclosed additional performance and pricing information on October 10. The model was post-trained on Qwen3.5-9B and is now available through Microsoft Foundry, with integration into OpenRouter as well.
This is not another general-purpose large model attempting to handle writing, search, programming, and complex reasoning all at once. Microsoft-Decision-1 has a narrow task: given a set of predefined options, select the more appropriate answer and assign probability scores to each candidate.
In other words, it is not responsible for explaining things clearly. It is responsible for deciding what should happen next.
In a typical Agent system, generating a response with a large model is only one part of the process. The system must also determine which model should handle a request, whether the result is acceptable, whether the task needs to be retried, whether a risky request should be blocked, and when it must be escalated to a human. In the past, these decisions were usually handled by rules, classifiers, or another general-purpose large model. Microsoft now hopes to use a model specifically trained for this layer.

35 Times Faster Than GPT-6 Sol, but That Does Not Simply Mean “35 Times More Capable”
Microsoft's internal evaluation covered 36 benchmarks and nearly 150,000 questions completely isolated from the training stage. According to the official results, Microsoft-Decision-1 ranked first in accuracy and was also the fastest-responding model among those evaluated:
- 4.5 times faster than Quyet-1.0-Large, which ranked second;
- 35 times faster than GPT-6 Sol;
- An average decision-flip probability of 1.3% in perturbation tests involving option rewording and changes in ordering;
- An officially reported decision-flip rate of zero in scenarios involving partial rewording or reversed ordering of option descriptions.
These figures need to be understood within their proper limits. Microsoft-Decision-1 performs constrained, structured decision-making, while general-purpose models such as GPT-6 Sol must handle open-ended generation, tool calling, and complex reasoning. The former only needs to score a limited set of candidates, whereas the latter may need to generate hundreds or even thousands of tokens. Comparing their latency directly is somewhat like comparing the time taken by a database index query with that of a complete data analysis report.
Therefore, “35 times faster” does not mean that its overall intelligence is 35 times greater. What it really shows is that for tasks such as classification, ranking, and routing, calling a full general-purpose model is often substantial overprovisioning.
That is where Microsoft-Decision-1 has its greatest value. It does not prove that specialized models can replace large language models. Instead, it reminds developers that not every step in a workflow is worth starting the most expensive and slowest model for.
Single-Pass Inference Scoring Targets High-Frequency Machine Decisions
Microsoft-Decision-1 is built on Qwen3.5-9B. Microsoft performed specialized post-training for decision-making tasks, enabling the model to score candidates through a single inference pass, rather than expanding into a lengthy reasoning process like a reasoning model and then parsing the conclusion from a natural-language answer.
This design is mainly suited to five types of tasks:
- Model routing: Send requests to a low-cost model, code model, or advanced reasoning model based on question difficulty, domain, and risk level.
- Content classification: Identify ticket topics, user intent, content risks, and feedback types.
- Task prioritization: Determine the processing priority of alerts, defects, review tasks, or customer requests.
- Result validation: Evaluate whether an upstream model's answer is acceptable and decide whether to approve it, retry it, or switch models.
- Workflow control: Choose between continuing execution, calling a tool, escalating for further processing, and handing the task to a human.
Microsoft's disclosed internal cases also largely centered on these scenarios. The Xbox Research team used it to classify more than 10,000 pieces of game feedback. The results were reportedly comparable in quality to GPT-6 Sol, while being more than 14 times faster and approximately 200 times cheaper. The Copilot team used it to evaluate AI-generated answers, with Microsoft stating that its performance was close to GPT-5.6 Luna.
These scenarios share several characteristics: high request volumes, limited output spaces, and individual decisions that do not justify consuming large numbers of generated tokens. Saving a few cents per invocation may not seem significant, but when the model sits at a routing or review node through which every request must pass, the cost difference is quickly amplified by the scale of usage.
$0.042 per Million Input Tokens, with Free Output
Microsoft-Decision-1 is priced at $0.042 per million input tokens, with output free of charge. Based on the exchange rate on October 10, this works out to approximately RMB 0.28 for the input cost.
“Free output” sounds aggressive, but it is related to the product's form factor. A decision model's output typically consists only of candidate labels, probability scores, and a small number of structured fields. It does not continuously generate long text like a generative model. Microsoft is effectively concentrating the main cost on the input side, making it easier for developers to estimate the expense of high-frequency usage.
For a routing service handling 10 million requests per day, the key issue is not simply that each invocation is cheap, but whether the model can reduce subsequent expensive calls. For example, Decision-1 can first assess the difficulty of a question, send simple requests to a small model, and route only a small number of complex requests to an advanced reasoning model. As long as routing accuracy is sufficiently high, the resulting savings may far exceed the cost of the decision model itself.
However, a low price cannot compensate for incorrect decisions. In scenarios such as incident response, content safety, and human review, a single routing error can cause losses far greater than the inference cost. Developers should pay closer attention to recall, false-positive rates, and probability calibration than to overall accuracy alone.
Probability Scores Are Useful, but Cannot Be Treated Directly as Confidence
Decision-1 assigns probabilities to candidate answers, which is more suitable for engineering systems than returning only a label. Applications can use these scores to set different thresholds: automatically execute high-confidence results, retry medium-confidence results, and send low-confidence or high-risk results to a human.
However, a model output of 0.9 does not inherently mean that there is a 90% probability of being correct in the real world. The reliability of a probability depends on how well it is calibrated against the data of a specific business. Ranking first on a training benchmark does not mean the model will maintain the same performance on private enterprise tickets, Chinese customer-service conversations, or highly specialized medical text.
Before formally integrating it into a production environment, at least the following validations are needed:
- Build an independent test set using real business data instead of relying solely on official benchmarks;
- Separately calculate precision, recall, and confusion matrices for each category;
- Check the actual accuracy within different confidence intervals to assess whether the probabilities are calibrated;
- Conduct perturbation tests involving option ordering, wording changes, missing fields, and malicious inputs;
- Set up rule-based fallbacks and human-escalation channels for high-risk decisions;
- Record the model version, candidates, scores, final action, and human correction results.
The design of the options deserves particular attention. A decision model can answer only from among the options provided by the developer. If the correct action should be “pause and collect more information,” but the candidates are limited to “approve” and “reject,” then even a highly accurate model can only choose between two incomplete answers. Many issues that appear to be model problems are actually caused by poorly designed decision spaces.
Safety Testing Covers Prompt Injection, but Production Environments Still Need Defense in Depth
Microsoft also evaluated the model using 11 safety benchmarks and 5,250 requests. The tests covered harmful content, jailbreak attacks, and prompt injection. The official conclusion was that Microsoft-Decision-1 could block malicious behavior while maintaining a high level of availability for normal requests.
Prompt injection is especially important for a model that may control an Agent's next action. An attacker does not necessarily need to make the model generate prohibited content. It may be enough to induce the decision layer to route a request to the wrong tool, skip a review, or approve a high-privilege operation, potentially compromising the entire workflow.
Therefore, Decision-1 is suitable as one part of a control system, but it should not be the sole security boundary. Permission checks, tool-parameter validation, usage quotas, deterministic rules, and audit logs should still be handled by systems outside the model. The model may recommend “allow execution,” but the caller's identity and resource permissions must still be checked before execution.
Foundry and OpenRouter Lower the Barrier to Experimentation; the Real Test Is Replaceability
Microsoft-Decision-1 is now available on Microsoft Foundry and has also been integrated into OpenRouter. For developers, this means they can connect the model to existing workflows without deploying Qwen3.5-9B themselves or maintaining a dedicated inference service.
OpenRouter's value lies in enabling quick comparisons across models, while Foundry is better suited to enterprises already using Microsoft's cloud, permission systems, and monitoring tools. Since regional availability, model identifiers, and quotas may vary across channels, actual availability should still be determined by what is displayed in the console.
From an architectural perspective, developers should not bind their business logic to a particular model name. A more robust approach is to define a standardized decision interface: provide task context and candidate actions as input, and return the candidates, scores, model version, and reason for rejection. This way, even if the underlying model is later replaced with a privately deployed classifier or a simple rules engine, the entire workflow does not need to be rewritten.
Microsoft has stated that it will later transfer this decision-training approach to other underlying models, including Microsoft AI (MAI) and OpenAI models. This suggests that Decision-1 is more like the starting point of a product line than an isolated release. Qwen3.5-9B is merely the current vehicle; what Microsoft truly wants to establish is a decision-making capability that can be replicated across different underlying models.
Assessment: It Is Not a Smaller Chat Model, but a Cheaper Control Layer
The direction of Microsoft-Decision-1 is sound. As Agents move from demos into production, the bottleneck is shifting from “Can the system generate an answer?” to “Can it reliably decide what to do next?” Using a general-purpose large model for every classification, validation, and routing task is not only expensive, but also makes latency and output formats more difficult to control.
By narrowing the problem, a specialized decision model makes speed, cost, and stability easier to optimize. The price of $0.042 per million input tokens is also low enough for high-frequency request paths. For teams that need multi-model routing, customer-service ticket classification, content moderation, or Agent workflow control, it has more practical value than adding another general-purpose chat model.
However, the core performance conclusions Microsoft has disclosed so far are based primarily on its own benchmarks and internal cases, with insufficient third-party replication. Ranking first across 36 tests and being 35 times faster than GPT-6 Sol can demonstrate that it is strong under specific task definitions, but cannot yet prove that it is more reliable across all enterprise decision-making scenarios.
In the short term, the most reasonable approach is not to hand over critical decisions to it all at once, but to first use it for low-risk routing and assisted scoring, observe it through shadow traffic for a period, and then gradually enable automatic execution. For developers, the most important question about Decision-1 is not its leaderboard position, but whether it can consistently reduce expensive model calls on real-world data without increasing erroneous escalations, missed blocks, or manual rework.
If both conditions can be met, Microsoft's release may be more than just an inexpensive classifier. It could become a genuinely useful layer of Agent infrastructure.
Sources
- IT Home: Microsoft launches the Microsoft-Decision-1 decision model, based on Qwen3.5-9B - Summarizes the model's release date, benchmark results, underlying model, availability channels, safety evaluations, and pricing information.



