Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentTutorialsPricingBlogContact

Microsoft-Decision-1: Price, Benchmarks, vs Jev

At a glanceQuick answers
What is Microsoft-Decision-1?
A decision model from Microsoft, released October 9, 2026, that reads an input and returns a probability for each of a fixed set of answers (yes or no, multiple choice, a rating or a rubric grade) instead of generating text.
What does it cost?
$0.042 per million input tokens, output tokens free, per Microsoft’s launch post and its OpenRouter listing. That is the same as Jev.
Is it better than Jev?
On Microsoft’s own table it is 1.2 points more accurate and about 2.8 times faster, while Jev stays ahead on calibration. No outside benchmark has checked it yet.
Editorial illustration on a near-white ground: a railway switch splits one track into three, ending in cards labelled ROUTE, ESCALATE and APPROVE with probability bars 0.91, 0.06 and 0.03, beside the figures 85 ms and $0.042
Fig 0One input, three fixed options, a probability for each (illustrative values). Made by CellCog's image agent, running GPT Image 2.5.

Microsoft released Microsoft-Decision-1 on October 9, 2026: a decision model that scores a fixed set of answers instead of writing text, priced at $0.042 per million input tokens with output free. That is exactly what TypeSafe charges for Jev, the model that started the category four weeks ago. Achint Srivastava, VP of software engineering in Microsoft’s Office of the CTO, published the launch post, and Satya Nadella posted it on X at 18:37 UTC. The opening line: “Decision models are quickly emerging as an important new category in AI.”

On this page · 8 sectionsOpen
  1. What Microsoft shipped
  2. Microsoft’s benchmark table
  3. What Microsoft says its own teams found
  4. The base model, and the price war
  5. What is not verified
  6. Where decision models fit for an AI employee
  7. What we are watching
  8. Sources
Key points6 · 7 min full read
  1. Three option cards with probability bars, one tall amber bar: a model picking among fixed options.
    Microsoft released Microsoft-Decision-1 on October 9, 2026: a decision model that returns a calibrated probability for each fixed option instead of writing text, in public preview on Microsoft Foundry and listed on OpenRouter.
  2. A stack of blocks with the top block being swapped for a new one: a base model to be replaced.
    It is Alibaba’s open-weight Qwen3.5-9B, post-trained by Microsoft; Microsoft says later versions will be rebased on its own MAI models and on OpenAI’s.
  3. Two equal price tags with an equals sign: the same price as Jev.
    The price is $0.042 per million input tokens with output free, the same as TypeSafe’s Jev. Cloudflare cut Clef-flash to $0.038 the same day.
  4. A stopwatch beside bars where the shortest bar is amber: the fastest median latency in Microsoft's table.
    In Microsoft’s own 36-benchmark comparison (147,137 questions) it averages 83.5% accuracy against Jev’s 82.3%, at an 85 ms median latency against Jev’s 240 ms.
  5. A dial gauge with its needle near the top: calibration close to perfect, but not first.
    On calibration it ranks third, 92.2 against Jev’s 93.7 and Quyet-1.0-Large’s 93.1. Microsoft added the Jev rows after it first published the post.
  6. A magnifying glass over a bar chart: the numbers still need an outside check.
    Every number is Microsoft’s own test, and its two launch posts give different figures for the same Xbox project. Treat the table as a claim until an outside benchmark repeats it.

§ 01What Microsoft shipped

Item Microsoft-Decision-1
Released October 9, 2026; public preview in Microsoft Foundry
Base model Qwen3.5-9B, post-trained by Microsoft for single-pass decision scoring
Answers Yes or no, multiple choice, ratings, and rubric grading of AI responses and agent actions
Output A calibrated probability for each fixed option
Price $0.042 per million input tokens; output tokens free
Where Microsoft Foundry; OpenRouter as microsoft/microsoft-decision-1, served by Azure
Context 32,768 tokens, per the OpenRouter listing
Input Text, per the OpenRouter listing
Table 1Microsoft-Decision-1 at launch (read 13:05 UTC October 10, 2026)

Microsoft names the jobs it is built for: routing, classification, prioritization, verification and workflow control. Its Foundry post calls it a model “for applications that need to choose among predefined options rather than generate open-ended text.”

§ 02Microsoft’s benchmark table

Microsoft ran its model and eight others across 36 benchmarks with 147,137 questions, which it says were kept blind from training. The table below is copied from the interactive chart in its post.

Model Accuracy Median latency Calibration
Microsoft-Decision-1 83.5% 85 ms (p95 125 ms) 92.2
Jev 1.13.0 (TypeSafe) 82.3% 240 ms 93.7
Quyet-1.0-Large 81.9% 380 ms 93.1
Surogate Rune 26B-A4B 79.7% 380 ms 91.8
GPT-6 Luna Decisions (OpenAI) 79.4% 300 ms 89.9
deck-31B 77.8% 400 ms 83.5
H2O-Lightning-4B v1.1 77.2% 210 ms 91.8
Strands-Decider 2B (AWS) 54.8%, on 23 of 36 benchmarks Not measured Not scored
GPT-6 Sol (reference) Not ranked 3,010 ms Not scored
Table 2Microsoft’s comparison: average accuracy over 36 benchmarks, median latency per request, calibration (100 = perfect)
Median latency per request in milliseconds, from Microsoft's tableBar chart of median latency from Microsoft's comparison: Microsoft-Decision-1 85 ms highlighted, H2O-Lightning-4B 210, Jev 240, GPT-6 Luna Decisions 300, Quyet-1.0-Large 380, Surogate Rune 380, deck-31B 400Microsoft-Decision-185H2O-Lightning-4B v1.1210Jev 1.13.0240GPT-6 Luna Decisions300Quyet-1.0-Large380Surogate Rune 26B-A4B380deck-31B400Median latency per request in milliseconds, from Microsoft's tableBar chart of median latency from Microsoft's comparison: Microsoft-Decision-1 85 ms highlighted, H2O-Lightning-4B 210, Jev 240, GPT-6 Luna Decisions 300, Quyet-1.0-Large 380, Surogate Rune 380, deck-31B 400Microsoft-Decision-185H2O-Lightning-4B v1.1210Jev 1.13.0240GPT-6 Luna Decisions300Quyet-1.0-Large380Surogate Rune 26B-A4B380deck-31B400
Fig 1Median latency per request in milliseconds, from Microsoft's table

Read the margins with the method in mind. The accuracy lead over Jev is 1.2 points. Jev is still better calibrated, which matters if you plan to act on the probability itself. The latency column is the JevBench v1.6.1 adjusted median, checked October 7, but Microsoft’s own number was measured through Foundry in the same region, so its model ran on home ground. On those figures it is 2.8 times faster than Jev and about 35 times faster than GPT-6 Sol.

Microsoft also tested whether the answer holds when the request is reworded: “Microsoft-Decision-1 changes its decision on 1.3% of perturbations on average with zero flips when option descriptions are paraphrased or when options are reversed or shuffled.” On safety, it reports 5,250 requests across 11 benchmarks covering harmful content, jailbreaks and prompt injection.

§ 03What Microsoft says its own teams found

  • Xbox Research sorted more than 10,000 pieces of player feedback into fixed themes. The launch post says quality was competitive with GPT-6 Sol at over 14 times the speed and 200 times less cost; its charts show 143 to 188 ms per text against 2.6 to 2.8 seconds for Sol, and about $11 against about $2,434 for a million texts. The Foundry post describes the same work as competitive with GPT-5 at 80 to 100 times faster. Both are Microsoft; the numbers differ.
  • Copilot used it to grade chat and agent responses and found it competitive with GPT-5.6 Luna at 100 times the speed.
  • Microsoft Discovery used it to grade experiments in an adaptive replanning loop; Microsoft says it was 46 times more consistent than an LLM grader at three times the speed.

§ 04The base model, and the price war

The model under the hood is Alibaba’s. In Microsoft’s words: “To build Microsoft-Decision-1, we post trained Qwen3.5-9B for fast, single-pass decision scoring and will soon rebase it on other models, including Microsoft AI (MAI) and OpenAI.” The Register led with exactly that, and counted the field: “All told, more than 100 such models are now vying for attention.”

Price is where the field is fighting. Per million input tokens, with output free or uncharged everywhere:

  • Cloudflare Clef-flash: $0.038, cut on October 9 in the same Cloudflare post that launched Clef-omni with audio and video input.
  • Perplexity Decisions API: $0.04.
  • Microsoft-Decision-1 and Jev: $0.042.
  • OpenAI Decisions API on gpt-6-luna: $0.10, per OpenAI’s guide.
  • Cloudflare Clef: $0.24.

So Microsoft did not undercut anyone; it matched the price that defines the category and competes on its own speed and accuracy figures. The earlier releases are compared in our Jev alternatives and OpenAI Decisions API pages.

§ 05What is not verified

  • Every number is Microsoft’s. No independent leaderboard has run Microsoft-Decision-1 yet; the JevBench positions in its chart describe the other models.
  • The table changed after launch. An editor’s note says the post was updated to add Jev’s accuracy and calibration.
  • Text only, for now. The OpenRouter listing shows text input; Microsoft’s computer-use demo does not settle what the API accepts.
  • The rebase. Microsoft gives no date for the MAI or OpenAI versions, and a new base would mean new numbers.

§ 06Where decision models fit for an AI employee

Our conflict, declared: every CellCog tier runs Claude Opus 5.5, we have not tested Microsoft-Decision-1, and we build AI employees that compete with agents some of these models are meant to steer.

The pitch for decision models is the small question asked thousands of times: which team gets this email, does this reply meet the rubric, should this step continue or stop. An employee answering those with a full language model pays for text nobody reads. A model that returns “route: 0.91” also hands you a threshold, so low-confidence cases can wait for a person. The caution is the same one Microsoft gives in its own post: check that the probabilities are calibrated on your data before you automate on them.

§ 07What we are watching

  • An outside run. The first independent JevBench or Vals result for Microsoft-Decision-1.
  • The rebase. A version on MAI or OpenAI models, and whether the price holds.
  • Inputs. Image or audio input on the API, now that Clef-omni takes both.

§ 08Sources

Frequently asked5 questions

Q1What is Microsoft-Decision-1 built on?

Microsoft post-trained Qwen3.5-9B, the 9-billion-parameter open-weight model from Alibaba’s Qwen team. Microsoft says it will rebase later versions on Microsoft AI (MAI) models and OpenAI models.

Q2Where can I use it?

In Microsoft Foundry, where it is in public preview, and on OpenRouter as microsoft/microsoft-decision-1, served by Azure with a 32,768-token context. The OpenRouter listing shows text input only.

Q3How does its price compare with other decision models?

It matches Jev at $0.042 per million input tokens. Perplexity’s Decisions API lists $0.04, Cloudflare’s Clef-flash $0.038 since October 9, Clef $0.24, and OpenAI’s Decisions API $0.10 per million input tokens.

Q4Can I trust the benchmark table?

It is Microsoft’s own run. The latency column uses the JevBench v1.6.1 method, but Microsoft measured its own model through Foundry, and it added the Jev rows after first publishing. Test it on your own decisions before relying on its probabilities.

Q5Does CellCog use Microsoft-Decision-1?

No. Every CellCog tier runs Claude Opus 5.5, and we have not tested Microsoft-Decision-1. This page reports what Microsoft published and what others have measured.

Published 10 October 2026 All Choosing a platform →