Microsoft released Microsoft-Decision-1 on October 9, 2026: a decision model that scores a fixed set of answers instead of writing text, priced at $0.042 per million input tokens with output free. That is exactly what TypeSafe charges for Jev, the model that started the category four weeks ago. Achint Srivastava, VP of software engineering in Microsoft’s Office of the CTO, published the launch post, and Satya Nadella posted it on X at 18:37 UTC. The opening line: “Decision models are quickly emerging as an important new category in AI.”
On this page · 8 sectionsOpen
Microsoft released Microsoft-Decision-1 on October 9, 2026: a decision model that returns a calibrated probability for each fixed option instead of writing text, in public preview on Microsoft Foundry and listed on OpenRouter.
It is Alibaba’s open-weight Qwen3.5-9B, post-trained by Microsoft; Microsoft says later versions will be rebased on its own MAI models and on OpenAI’s.
The price is $0.042 per million input tokens with output free, the same as TypeSafe’s Jev. Cloudflare cut Clef-flash to $0.038 the same day.
In Microsoft’s own 36-benchmark comparison (147,137 questions) it averages 83.5% accuracy against Jev’s 82.3%, at an 85 ms median latency against Jev’s 240 ms.
On calibration it ranks third, 92.2 against Jev’s 93.7 and Quyet-1.0-Large’s 93.1. Microsoft added the Jev rows after it first published the post.
Every number is Microsoft’s own test, and its two launch posts give different figures for the same Xbox project. Treat the table as a claim until an outside benchmark repeats it.
§ 01What Microsoft shipped
| Item | Microsoft-Decision-1 |
|---|---|
| Released | October 9, 2026; public preview in Microsoft Foundry |
| Base model | Qwen3.5-9B, post-trained by Microsoft for single-pass decision scoring |
| Answers | Yes or no, multiple choice, ratings, and rubric grading of AI responses and agent actions |
| Output | A calibrated probability for each fixed option |
| Price | $0.042 per million input tokens; output tokens free |
| Where | Microsoft Foundry; OpenRouter as microsoft/microsoft-decision-1, served by Azure |
| Context | 32,768 tokens, per the OpenRouter listing |
| Input | Text, per the OpenRouter listing |
Microsoft names the jobs it is built for: routing, classification, prioritization, verification and workflow control. Its Foundry post calls it a model “for applications that need to choose among predefined options rather than generate open-ended text.”
§ 02Microsoft’s benchmark table
Microsoft ran its model and eight others across 36 benchmarks with 147,137 questions, which it says were kept blind from training. The table below is copied from the interactive chart in its post.
| Model | Accuracy | Median latency | Calibration |
|---|---|---|---|
| Microsoft-Decision-1 | 83.5% | 85 ms (p95 125 ms) | 92.2 |
| Jev 1.13.0 (TypeSafe) | 82.3% | 240 ms | 93.7 |
| Quyet-1.0-Large | 81.9% | 380 ms | 93.1 |
| Surogate Rune 26B-A4B | 79.7% | 380 ms | 91.8 |
| GPT-6 Luna Decisions (OpenAI) | 79.4% | 300 ms | 89.9 |
| deck-31B | 77.8% | 400 ms | 83.5 |
| H2O-Lightning-4B v1.1 | 77.2% | 210 ms | 91.8 |
| Strands-Decider 2B (AWS) | 54.8%, on 23 of 36 benchmarks | Not measured | Not scored |
| GPT-6 Sol (reference) | Not ranked | 3,010 ms | Not scored |
Read the margins with the method in mind. The accuracy lead over Jev is 1.2 points. Jev is still better calibrated, which matters if you plan to act on the probability itself. The latency column is the JevBench v1.6.1 adjusted median, checked October 7, but Microsoft’s own number was measured through Foundry in the same region, so its model ran on home ground. On those figures it is 2.8 times faster than Jev and about 35 times faster than GPT-6 Sol.
Microsoft also tested whether the answer holds when the request is reworded: “Microsoft-Decision-1 changes its decision on 1.3% of perturbations on average with zero flips when option descriptions are paraphrased or when options are reversed or shuffled.” On safety, it reports 5,250 requests across 11 benchmarks covering harmful content, jailbreaks and prompt injection.
§ 03What Microsoft says its own teams found
- Xbox Research sorted more than 10,000 pieces of player feedback into fixed themes. The launch post says quality was competitive with GPT-6 Sol at over 14 times the speed and 200 times less cost; its charts show 143 to 188 ms per text against 2.6 to 2.8 seconds for Sol, and about $11 against about $2,434 for a million texts. The Foundry post describes the same work as competitive with GPT-5 at 80 to 100 times faster. Both are Microsoft; the numbers differ.
- Copilot used it to grade chat and agent responses and found it competitive with GPT-5.6 Luna at 100 times the speed.
- Microsoft Discovery used it to grade experiments in an adaptive replanning loop; Microsoft says it was 46 times more consistent than an LLM grader at three times the speed.
§ 04The base model, and the price war
The model under the hood is Alibaba’s. In Microsoft’s words: “To build Microsoft-Decision-1, we post trained Qwen3.5-9B for fast, single-pass decision scoring and will soon rebase it on other models, including Microsoft AI (MAI) and OpenAI.” The Register led with exactly that, and counted the field: “All told, more than 100 such models are now vying for attention.”
Price is where the field is fighting. Per million input tokens, with output free or uncharged everywhere:
- Cloudflare Clef-flash: $0.038, cut on October 9 in the same Cloudflare post that launched Clef-omni with audio and video input.
- Perplexity Decisions API: $0.04.
- Microsoft-Decision-1 and Jev: $0.042.
- OpenAI Decisions API on gpt-6-luna: $0.10, per OpenAI’s guide.
- Cloudflare Clef: $0.24.
So Microsoft did not undercut anyone; it matched the price that defines the category and competes on its own speed and accuracy figures. The earlier releases are compared in our Jev alternatives and OpenAI Decisions API pages.
§ 05What is not verified
- Every number is Microsoft’s. No independent leaderboard has run Microsoft-Decision-1 yet; the JevBench positions in its chart describe the other models.
- The table changed after launch. An editor’s note says the post was updated to add Jev’s accuracy and calibration.
- Text only, for now. The OpenRouter listing shows text input; Microsoft’s computer-use demo does not settle what the API accepts.
- The rebase. Microsoft gives no date for the MAI or OpenAI versions, and a new base would mean new numbers.
§ 06Where decision models fit for an AI employee
Our conflict, declared: every CellCog tier runs Claude Opus 5.5, we have not tested Microsoft-Decision-1, and we build AI employees that compete with agents some of these models are meant to steer.
The pitch for decision models is the small question asked thousands of times: which team gets this email, does this reply meet the rubric, should this step continue or stop. An employee answering those with a full language model pays for text nobody reads. A model that returns “route: 0.91” also hands you a threshold, so low-confidence cases can wait for a person. The caution is the same one Microsoft gives in its own post: check that the probabilities are calibrated on your data before you automate on them.
§ 07What we are watching
- An outside run. The first independent JevBench or Vals result for Microsoft-Decision-1.
- The rebase. A version on MAI or OpenAI models, and whether the price holds.
- Inputs. Image or audio input on the API, now that Clef-omni takes both.
§ 08Sources
- Achint Srivastava, Microsoft, Introducing Microsoft-Decision-1, our model for fast decision-making, October 9, 2026 (updated with Jev rows), read 13:05 UTC October 10.
- Saumil Shrivastava, Microsoft Foundry Blog, Introducing Microsoft-Decision-1 in Microsoft Foundry, October 9, 2026.
- Satya Nadella, post on X, October 9, 2026, 18:37 UTC.
- OpenRouter, microsoft/microsoft-decision-1 endpoint listing, read 13:05 UTC October 10.
- Thomas Claburn, The Register, Microsoft leans on open weight model from Chinese AI lab to challenge Jev, October 10, 2026.
- Cloudflare, Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash, October 9, 2026.
Q1What is Microsoft-Decision-1 built on?
Microsoft post-trained Qwen3.5-9B, the 9-billion-parameter open-weight model from Alibaba’s Qwen team. Microsoft says it will rebase later versions on Microsoft AI (MAI) models and OpenAI models.
Q2Where can I use it?
In Microsoft Foundry, where it is in public preview, and on OpenRouter as microsoft/microsoft-decision-1, served by Azure with a 32,768-token context. The OpenRouter listing shows text input only.
Q3How does its price compare with other decision models?
It matches Jev at $0.042 per million input tokens. Perplexity’s Decisions API lists $0.04, Cloudflare’s Clef-flash $0.038 since October 9, Clef $0.24, and OpenAI’s Decisions API $0.10 per million input tokens.
Q4Can I trust the benchmark table?
It is Microsoft’s own run. The latency column uses the JevBench v1.6.1 method, but Microsoft measured its own model through Foundry, and it added the Jev rows after first publishing. Test it on your own decisions before relying on its probabilities.
Q5Does CellCog use Microsoft-Decision-1?
No. Every CellCog tier runs Claude Opus 5.5, and we have not tested Microsoft-Decision-1. This page reports what Microsoft published and what others have measured.
