Microsoft Decision-1 AI model launch
TECH

Microsoft Decision-1 AI model launch

31+
Signals

Strategic Overview

  • 01.
    Microsoft launched Microsoft-Decision-1 in Microsoft Foundry, a model that reads supplied content plus a set of predefined answer options and returns a calibrated probability score for each one instead of generating open-ended text.
  • 02.
    Decision-1 is built on Alibaba's open-weight 9-billion-parameter Qwen3.5-9B base model, post-trained by Microsoft engineers into a closed product with no published weights.
  • 03.
    Pricing is $0.042 per million input tokens with output tokens free, because the model returns only scores rather than generated prose.
  • 04.
    Microsoft says Decision-1 is the most accurate model it tested across 36 benchmarks covering nearly 150,000 questions, running 2.5 times faster than runner-up H2O-Lightning-4B and roughly 35 times faster than GPT-6 Sol.

Microsoft's Newest Commercial Model Is Alibaba's Model Underneath

The most consequential detail of this launch is not the speed claim - it is the provenance. Microsoft did not build a new foundation model for structured decision-making, and it did not take one from its OpenAI partnership either. It post-trained Alibaba's open-weight Qwen3.5-9B [1], a 9-billion-parameter base model, into a product its own CEO personally unveiled [15]. For a company that spends at foundation-model scale, choosing a Chinese lab's open release as the substrate for a flagship Foundry SKU is a statement about where value now sits: not in the weights, but in the post-training, the calibration, and the distribution.

The direction of travel then reverses. Decision-1 ships with no open weights at all, and the Hugging Face listing for it returns a 404, which means self-hosting and fine-tuning are not options for customers [3]. Microsoft has also said it plans to rebase the model on its own MAI models and on OpenAI models in addition to the current Qwen base [4], without publishing a timeline or a backward-compatibility guarantee [5]. That is a real integration risk for anyone wiring Decision-1 into agent control loops today, because the thing being swapped out is precisely the thing teams will have tuned their thresholds against: the scoring distribution and the latency profile of a specific 9B model.

Developers noticed the asymmetry immediately. On Hacker News, commenters pointed out that Decision-1 is based on one of the smaller Qwen models, just like Cloudflare's Clef, Strands decider, and a plethora of others released [6], which deflates the premise that Microsoft built something structurally proprietary. The sharpest version of the complaint circulating in developer communities was simpler and more cynical: that the impressive part is taking something open-weight and making it closed.

By The Numbers: The Comparison Microsoft Chose Not To Run

By The Numbers: The Comparison Microsoft Chose Not To Run
Price per million input tokens across the decision-model field, where Decision-1 ties Jev and sits above Clef-flash.

Decision-1's pricing is genuinely aggressive in absolute terms. Input tokens cost $0.042 per million and output tokens are free, since the model emits scores rather than prose [7]. The accuracy and latency claims are stacked just as carefully: Microsoft reports best-in-test accuracy of 83.5% across 36 benchmarks covering roughly 150,000 questions at an 85-millisecond median latency [8], about 2.5 times faster than runner-up H2O-Lightning-4B [9], and the headline number everyone repeated was 35 times faster than GPT-6 Sol [2].

That last comparison is the one worth interrogating, because GPT-6 Sol is a general-purpose generative model, not a decision model. Compared against the set a buyer would actually shortlist, the picture flattens: Decision-1 matches TypeSafe's Jev to the cent at $0.042 per million input tokens, sits above Cloudflare's Clef-flash at $0.038 per million tokens with Apache 2.0 weights you can host yourself, undercuts OpenAI's Decisions API at roughly $0.10 per million by about 2.4 times, and comes in far below Cloudflare's larger Clef at $0.24 per million [10]. The eesel.ai analysis puts the objection bluntly: Microsoft only compares the price of theirs to GPT Sol, not GPT Terra, or GPT Luna, and certainly not Jev [10].

Where Microsoft does hold a measurable edge is context length - a 32,768-token window against Clef-flash's 24,576 [10]- which matters for decisions that require reading a long ticket thread, a diff, or an incident timeline before scoring options. That is a narrower and more defensible pitch than 35x, and notably not the one the launch led with.

Calibrated Is Not The Same As Correct

Mechanically, Decision-1 is doing something unusual enough to be interesting. Instead of asking a model to write an answer and then parsing it, you hand it content plus a fixed list of candidate options, and it returns a calibrated probability for each one [11]. The robustness numbers Microsoft publishes are the strongest technical part of the release: the model changes its decision on only 1.3% of input perturbations on average, with zero flips when option descriptions are paraphrased or when the options are reversed or shuffled [7]. Anyone who has fought position bias in LLM-as-judge pipelines will recognize why that is the headline a practitioner should care about. Remio.ai's read is that this separation of structured choices from open-ended generation is the actual story of the launch [12], not the executive endorsement attached to it.

The risk arrives with the use cases. Microsoft markets Decision-1 for routing, classification, prioritization, verification, workflow and agent control, model routing, data labeling, AI judging, incident response routing, content filtering, safety screening, computer UI use and robotics [7]. Those are exactly the places where a confident number is more dangerous than hedged prose. The Decoder's analysis makes the point directly: a calibrated confidence score is not a correctness guarantee, benchmark accuracy is not safety for high-impact workflows, and the published results still want independent verification [8]. A generative model that waffles at least signals doubt to whatever reads it next; a score of 0.91 handed to an autonomous agent reads as permission.

Practitioner reaction has already found the seam between the datasheet and deployment. In machine-learning communities, one hands-on tester reported production latency in Foundry several times higher than the advertised median, along with request-rate limits that would reshape any high-volume routing design. Others questioned the comparison methodology itself, noting that a competitor's own model card reports a far lower measured median than the figure used in Microsoft's chart - a discrepancy that suggests Microsoft timed rivals differently than it timed itself. Nothing there contradicts the model's quality; it does argue for benchmarking it on your own traffic before trusting the marketing curve.

A Market Category That Formed In About Thirty Days

Step back from Microsoft and the more durable story is category formation at unusual speed. TypeSafe AI's Jev opened the decision model or System One category in September 2026 [13]. OpenAI followed on September 29 with a Decisions API built on a specialized GPT-6 Luna that classifies in roughly 150 milliseconds against about 1.6 seconds for a standard Luna call [13]. Cloudflare answered in October with Clef and Clef-flash, Apache 2.0 and Jev-API compatible so that switching costs approach zero [13]. Microsoft arrived on October 9 into a field that already contains more than 100 decision models [14].

Why every major platform wants this specific primitive comes down to unit economics. Microsoft's own internal example is the tell: Xbox Research classified more than 10,000 player feedback items with quality competitive with GPT-5, while running between 14 and 100 times faster at roughly 200 times lower cost [2][15]. Any company running high-volume classification, moderation, or agent routing through a frontier LLM is overpaying by a similar multiple, and whoever owns that layer owns a metered toll on every agent decision. Coverage of the launch positions Decision-1 inside Microsoft's four-layer stack of Azure, Fabric, Foundry and Copilot [15].

The catch is that a primitive this easy to replicate resists moats. With an open-weight base, a published scoring interface, and Jev-compatible APIs already standard, differentiation collapses toward price, latency and distribution - and Microsoft only wins the third outright. That split is visible in how the launch was received: developer-facing commentary skewed toward mechanism and replication, with a developer video on the release framing it as a distinctly odd move, while a separate finance-adjacent audience treated the announcement as a stock catalyst rather than a product release. Both readings can be correct. Decision-1 is a modest technical step and a significant distribution event, and only one of those is hard to copy.

Historical Context

2026-09
TypeSafe AI launched Jev, the first model in the decision model or System One category, creating the category Decision-1 now enters.
2026-09-29
OpenAI unveiled its Decisions API, built on a specialized version of GPT-6 Luna that returns classified answers in roughly 150 milliseconds versus about 1.6 seconds for a standard Luna call.
2026-10
Cloudflare released Clef on Qwen3.8-27B and Clef-flash on Qwen3.5-9B under Apache 2.0, fully Jev-API compatible and self-hostable, as open-weight rivals to the closed decision-model products.
2026-10-09
Microsoft launched Microsoft-Decision-1 in Microsoft Foundry, joining a field that already contains more than 100 decision models.

Power Map

Key Players
Subject

Microsoft Decision-1 AI model launch

MI

Microsoft

Developer and seller of Decision-1 through Microsoft Foundry, unveiled personally by CEO Satya Nadella and positioned inside Microsoft's four-layer AI stack of Azure, Fabric, Foundry and Copilot as a bid to own the decision-scoring middleware layer.

AL

Alibaba (Qwen team)

Supplied the open-weight Qwen3.5-9B base model that Microsoft post-trained into Decision-1, which leaves Microsoft's newest commercial decision model dependent on a rival's foundation until it rebases on MAI or OpenAI models.

TY

TypeSafe AI (Jev)

First mover in the decision model or System One category, and now the direct pricing benchmark - Decision-1 matches Jev to the cent at $0.042 per million input tokens.

CL

Cloudflare (Clef and Clef-flash)

Shipped Apache 2.0 open-weight competitors built on Qwen3.8-27B and Qwen3.5-9B that are API-compatible with Jev; Clef-flash undercuts Decision-1 at $0.038 per million tokens and is self-hostable, which caps how much Microsoft can charge for the same workload.

OP

OpenAI

Launched a competing Decisions API built on GPT-6 Luna on September 29, 2026; Decision-1 undercuts that API on price even though Microsoft says it intends to eventually rebase Decision-1 on OpenAI models.

XB

Xbox Research (Microsoft internal)

Early internal adopter that used Decision-1 to classify more than 10,000 player feedback items, reporting GPT-5-competitive quality at far lower latency and cost - the proof point Microsoft leans on publicly.

Fact Check

15 cited
  1. [1] Microsoft Skipped OpenAI's Decision Model and Built Its Own on Alibaba's Qwen
  2. [2] Microsoft Decision-1 AI Model
  3. [3] Microsoft Decision-1
  4. [4] Microsoft Launches Decision-1: Fast Decision-Scoring Model Built on Qwen3.5-9B
  5. [5] Microsoft Decision-1 AI Model: 35x Speed
  6. [6] Microsoft-Decision-1 Hacker News discussion thread
  7. [7] Microsoft-Decision-1 Model in Microsoft Foundry
  8. [8] Microsoft's Decision-1 Model Enters the Fast-Growing AI Decision Model Race
  9. [9] Microsoft Enters Decision Model Race With Qwen-Based Microsoft-Decision-1
  10. [10] Microsoft Decision-1 Pricing
  11. [11] Microsoft Decision-1: A Cheap Text-Only Decision Model Built on Qwen
  12. [12] Microsoft Decision-1 Model Arrives but Nadella's Endorsement Is Not the Main Story
  13. [13] Decision Models
  14. [14] Microsoft Launches Decision-1 as 100 AI Models Enter the Market
  15. [15] Microsoft's Nadella Unveils Tech Giant's New Fast Decision-Making AI Model

Source Articles

Top 3

THE SIGNAL.

Analysts

“Argues Decision-1's benchmark claims still need independent verification, that calibrated confidence is not the same as correctness, and that benchmark performance does not equal safety in high-impact workflows that still need human oversight. Its summary: this launch changes the available architecture more than it changes what AI can understand.”

The Decoder
Tech news and analysis outlet, the-decoder.com

“Calls Microsoft's framing misleading because the comparison is run against GPT-6 Sol rather than the actual competitive set: Microsoft only compares the price of theirs to GPT Sol, not GPT Terra, or GPT Luna, and certainly not Jev. Against that set, Decision-1 ties Jev on price and is more expensive than Cloudflare's self-hostable Clef-flash.”

eesel.ai
Independent AI pricing and product analysis site

“Pushed back on both the pricing and the novelty claims, noting it is price competitive but at $0.042 per million input tokens that is the same as Jev rather than cheaper, and that this is based on one of the smaller Qwen models, just like Cloudflare's Clef, Strands decider, and a plethora of others released.”

Hacker News commenters
news.ycombinator.com discussion thread on the Decision-1 announcement

“Argues the substance of the launch is the architectural separation of structured choices from open-ended generation, not Nadella's personal endorsement of the model.”

Remio.ai
Independent AI commentary site
The Crowd

“Introducing Microsoft-Decision-1, our new model for fast decision-making. It delivers top performance on structured decision tasks, outperforming both LLMs and other decision models in latency and quality. We're already testing it across Microsoft for everything from incident response to...”

@@satyanadella11325

“OpenAI Decision crushed by @Microsoft-Decision-1 at Space Invaders 👾 Decision-1 decided 3.5x faster so it scored 3,320 and cleared 4 waves while OpenAI's decision API scored 1,300 taking over 300ms per decision on average Run decision models -> https://t.co/RbcCOIht9R”

@@atomic_chat_hq783

“Microsoft-Decision-1 is live on OpenRouter. Microsoft's numbers: highest accuracy across 36 blind benchmarks (~150K questions), 4.5x faster than the runner-up, 35x faster than GPT-6 Sol, and decisions flip on just 1.3% of perturbed inputs. Post-trained from Qwen3.5-9B. $0.042/M”

@@OpenRouter594

“Introducing Microsoft-Decision-1, our model for fast decision-making”

@u/Greedyanda572
Broadcast
Microsoft-Decision-1 (What an odd release?!?)

Microsoft-Decision-1 (What an odd release?!?)

Microsoft Decision-1: Same Price as Jev, Built on Qwen

Microsoft Decision-1: Same Price as Jev, Built on Qwen

Microsoft's New Decision Model: A Fresh AI Catalyst and the $537 Test | Oct 9

Microsoft's New Decision Model: A Fresh AI Catalyst and the $537 Test | Oct 9