Microsoft AI Releases Microsoft-Decision-1: A Qwen3.5-9B Decision-Scoring Model
Oct 9, 2026, 11:03 PM · MarkTechPost
Microsoft is betting the next useful AI layer does not talk at all. Decision-1 picks from fixed options fast and cheap, but every number behind it comes from Microsoft's own scoreboard.
Why it matters
Microsoft has released Microsoft-Decision-1, a model that does not generate text. It reads an input and a fixed set of answer options and returns a calibrated probability for each one. It is post-trained from Alibaba's Qwen3.5-9B and available as a hosted API in Microsoft Foundry and on OpenRouter, according to MarkTechPost.
The use cases are the unglamorous plumbing of AI products: routing requests, classifying items, checking whether an answer is grounded in evidence, grading other models' outputs and deciding whether an agent's proposed action should go ahead. Those calls happen constantly inside software, so speed and price matter more than eloquence. Microsoft reports an 85 millisecond median latency and pricing of $0.042 per million input tokens, with output free.
From the desk
We like the idea here more than the packaging. A lot of what people call AI in production is really a chain of small judgments: is this ticket billing or technical, is this answer supported by the document, should this agent be allowed to send that email. Using a large chat model for each of those is slow and expensive, and the free-text output has to be parsed back into something software can act on. A model that just returns a probability per option, with explicit abstain choices like cannot tell, is a cleaner fit.
The agent-control angle is the most interesting to us. Microsoft lists grading proposed agent actions as a supported format. A fast, cheap checker that sits in front of an agent and scores whether an action is in bounds is exactly the kind of guardrail the industry needs more of right now. If it works as described, it makes oversight cheaper to run at every step rather than occasionally.
Now the caveats, and they are real. This is closed weights, hosted only, with the parameter count undisclosed. MarkTechPost notes every benchmark is vendor-run. Microsoft's 36-benchmark comparison puts the model at 83.5% average accuracy, ahead of Quyet-1.0-Large at 81.9%, but the latency comparison is disputed. Microsoft timed its own model through Foundry while competitor figures use JevBench's adjusted median, and H2O.ai's model card says that adjustment inflates its model's 29 millisecond measured median to 210. Microsoft-Decision-1 does not appear on the JevBench board at all. Until independent results land, we would treat the speed lead as unproven.
There is also a quieter tension. OpenRouter notes the weights update continually while the API shape stays the same. For a model meant to make consistent decisions, a moving target underneath a fixed interface could make behavior shift without anyone changing a line of code. Teams using it for anything they audit will want to pin and test.
We give Microsoft credit for an unusually explicit limits section. The model card says not to use it as the sole automated decision-maker on credit, employment, housing, healthcare or legal rights, and that applications must define thresholds, escalation paths and human oversight. That is the right warning. The risk is that a model built to output a confident probability gets wired into exactly those decisions anyway, because it is cheap and it does not explain itself. A decision with no rationale is hard to appeal.
Our read: a sensible product in a category that is clearly forming, with a strong price, held back by self-reported numbers. I'm watching for independent benchmarks and for which agent platforms adopt it as a guardrail.
Context
Decision models are an emerging category of systems that score a closed set of options instead of writing prose. MarkTechPost's comparison includes Quyet-1.0-Large, H2O.ai's H2O-Lightning-4B and OpenAI's GPT-6 Luna Decisions, which it lists at $0.10 per million input tokens. Microsoft says it plans to rebase future versions on MAI and OpenAI models.
Who feels it
- Developers
- A cheap, structured decision call can replace parsing chat output for routing and classification, but it runs on a separate Decisions API, not standard chat completions SDKs.
- Agent builders
- Scoring proposed agent actions before they run is a practical guardrail pattern this model is designed for.
- Compliance teams
- No rationales in the output and continually updating weights complicate audits; the model card rules out sole use in high-stakes decisions.
- Open-model vendors
- Competitors like H2O.ai are already disputing the latency comparison, and open-weight rivals can self-host where Microsoft cannot.
What to watch
- Independent JevBench or third-party results for Microsoft-Decision-1
- Whether Microsoft resolves the latency methodology dispute raised by H2O.ai
- Future versions rebased on MAI or OpenAI models, as Microsoft plans
- Whether Microsoft offers version pinning given continually updating weights
- Adoption inside agent frameworks as an action-approval layer
Companies: Microsoft