Microsoft Decision-1 model launch
TECH

Microsoft Decision-1 model launch

25+
Signals

Strategic Overview

  • 01.
    Microsoft launched Decision-1 on October 9, 2026, a non-generative decision-scoring model built by post-training Alibaba's open-weight Qwen3.5-9B, which returns calibrated probabilities for a fixed set of answer options instead of generating open-ended text.
  • 02.
    Pricing is set at $0.042 per million input tokens with output tokens free, available now via Microsoft Foundry with OpenRouter access coming soon, targeting routing, classification, agent action selection, and safety screening.
  • 03.
    Microsoft reports the highest accuracy across a 36-benchmark sweep of roughly 150,000 withheld questions and the fastest P50 latency among compared models, while internal Xbox Research and Copilot teams cite large speed and cost gains from early deployments.
  • 04.
    Independent analysis disputes Microsoft's headline speed claim, arguing it compares a raw Microsoft-measured figure against a formula-adjusted competitor number, with the real raw-timing advantage closer to 1.4x rather than the advertised 4.5x.

Deep Analysis

A Genuinely New Model Category, or Just an LLM With One Token?

Nadella personally unveiled Decision-1, framing it as beating both general-purpose LLMs and other decision models on latency and quality[1]. But the sharper reaction didn't come from the executive framing - it came from engineers picking apart what a "decision model" even is. In the launch thread on r/singularity, the debate split into two camps: one argued Decision-1 is "literally" nothing more than a standard LLM with its output truncated to a single token and constrained logprobs, while others countered that purpose-built decision-scoring models skip the autoregressive reasoning overhead entirely, making them structurally more efficient for narrow choice tasks rather than merely differently constrained. A related thread of skepticism questioned the market fit itself, noting that Copilot already wraps third-party chat models and framing Decision-1 as more of an Azure-ecosystem enterprise play than a frontier-chat competitor.

Independent commentary largely sided with the view that the category label matters less than the underlying architectural signal. "Decision-1 is notable because it challenges the default use of general-purpose LLMs, not because a CEO shared its launch," argued Sophie Larsen[2], who cautioned that Microsoft's internal benchmark numbers still await independent verification. A separate strand of the Reddit debate went further back than LLMs entirely, with commenters pointing out that classifier-style decision models predate the generative-AI era - BERT-era classifiers did something similar - and are, in their words, "beyond trivial" to build with modern tooling, which raises the question of whether Decision-1's real contribution is a new primitive or simply a well-packaged, well-priced version of an old one.

The Benchmarks Don't Actually Resolve the Competition

Microsoft's own comparison table against TypeSafe's Jev doesn't show a clean win. Decision-1 reports higher accuracy (83.5% vs 82.3%) and a far lower median latency (85ms vs 240ms), but Jev leads on calibration (93.7 vs 92.2)[3], meaning the two models are already trading wins across different metrics rather than one clearly beating the other.

The headline "4.5x faster than the runner-up" claim is more contested still. Independent analysis found the comparison pairs Microsoft's own 85ms Foundry-measured figure against a 380ms "adjusted" JevBench number derived from a formula - not a raw measurement - and that JevBench's own notes describe that adjustment as "an assumption, not a hardware normalization." On a true apples-to-apples raw-timing basis, the gap closes to roughly 1.4x (85ms vs 116ms), not 4.5x[4]. A parallel critique surfaced on r/machinelearningnews, where the post itself flagged that H2O.ai's own model card reportedly lists a raw median latency of just 29ms for H2O-Lightning-4B, far below the 210ms figure Microsoft cites for it in its chart. A separate thread on r/singularity raised a related objection specific to Jev: commenters noted that Jev's real strength is batched throughput (255 decisions in roughly 150ms) rather than single-request latency, making a single-decision-latency comparison against it potentially misleading on its own terms. Taken together, nearly every specific number in Microsoft's launch materials has drawn a methodology challenge from a different angle.

Not every independent check cut against Microsoft. One developer ran an informal head-to-head on X, pitting Decision-1 against OpenAI's own decision-scoring API in a reaction-speed game, and reported Decision-1 responding roughly 3.5 times faster and clearing more waves than OpenAI's model managed. It is a single informal test rather than a published benchmark, but unlike Microsoft's internal figures, it came from outside the company - a reminder that the dispute here is over methodology and framing, not over whether Decision-1 is fast in absolute terms.

Microsoft's Decision Layer Is Built on a Chinese Open-Weight Model

Decision-1 is post-trained from Alibaba's open-weight Qwen3.5-9B rather than a Microsoft or OpenAI base model[5]. That is a pragmatic bet: for a narrow decision-scoring task, base-model provenance matters less than speed and price, and post-training an existing open-weight model let Microsoft ship a priced, benchmarked product quickly rather than training a foundation model from scratch. One tech-explainer video summary noted Decision-1 launched at effectively the same input price as TypeSafe's Jev despite running on a third-party Chinese base model, underscoring how commoditized decision-scoring pricing has become across vendors regardless of what sits underneath each product.

The choice drew its own criticism. On r/machinelearningnews, one commenter objected to Microsoft serving a fine-tuned version of an open-weight Chinese model commercially without releasing the resulting fine-tuned weights back to the community, arguing it inverts the usual open-source exchange. A separate exchange challenged the broader decision-model funding picture: one commenter called a rival startup's roughly $800 million valuation "lunacy," arguing incumbents like Microsoft would simply out-execute smaller players once the category matured; another commenter pushed back - a live dispute that, either way, underscores how unsettled the question of who actually wins this niche still is.

The Real Bet Is Volume: Routing Cheap Decisions Away From Frontier LLMs

Microsoft's pricing - $0.042 per million input tokens with output free - is explicitly framed by analysts as a response to cost becoming the dominant factor in how developers architect agentic AI pipelines, where a single task can generate dozens of discrete routing or classification sub-decisions[6]. The strategic logic is to own a distinct "decision layer" of the AI stack across Azure, Fabric, Foundry, and Copilot, monetized separately from general-purpose LLM inference rather than competing for the same chat-completion dollar.

A Reddit commenter sketched the practical version of that pattern directly: pair a decision model with a frontier LLM so the LLM offloads simple yes/no or classification sub-decisions to the cheaper model, a substitution that "adds up" in savings at scale. Microsoft's own cited internal deployments look like early proof points for exactly that pattern - Xbox Research and the Copilot team both reported large speed and cost reductions on production workloads[1]- though those figures carry the same caveat as the rest of the launch: they come from Microsoft itself, with no independent benchmark yet published to confirm them.

Historical Context

2026-10-09
Microsoft announced and launched Microsoft-Decision-1 in Microsoft Foundry, with OpenRouter availability following shortly after.

Power Map

Key Players
Subject

Microsoft Decision-1 model launch

MI

Microsoft

Developer and publisher of Decision-1, positioning it as the decision layer of its agentic-AI stack across Azure, Fabric, Foundry, and Copilot, with pricing designed to capture high-volume, low-value decision calls separately from general LLM usage.

AL

Alibaba (Qwen)

Supplier of the open-weight Qwen3.5-9B base model that Microsoft post-trained to create Decision-1, rather than using a Microsoft or OpenAI foundation model.

TY

TypeSafe (Jev)

Direct competitor whose decision model beats Decision-1 on calibration (93.7 vs 92.2) even as Microsoft claims higher accuracy and lower latency, meaning the competitive outcome splits by metric rather than resolving outright.

SA

Satya Nadella (Microsoft CEO)

Personally unveiled and endorsed the launch, publicly framing Decision-1 as outperforming both general LLMs and rival decision models, lending executive weight to claims that outside analysts later contested.

OP

OpenRouter

Third-party model marketplace extending Decision-1's distribution beyond the Azure ecosystem, listed as 'coming soon' alongside the immediate Foundry launch.

Fact Check

6 cited
  1. [1] Microsoft Launches Decision-1 Model in Foundry
  2. [2] Microsoft Decision-1 Model Arrives, But Nadella's Endorsement Is Not the Main Story
  3. [3] Microsoft Decision-1
  4. [4] Microsoft Decision-1: 4.5x Faster Than Runner-Up? Raw Timing Gives 1.4x
  5. [5] Microsoft-Decision-1: Decision Model in Foundry
  6. [6] Microsoft Launches Decision-1 at $0.042 Per Million Tokens

Source Articles

Top 5

THE SIGNAL.

Analysts

“Argues Microsoft's '4.5x faster than the runner-up' claim conflates a raw Microsoft-measured figure with a formula-adjusted competitor figure, and that the real raw-timing lead is only about 1.4x.”

P.K. Sharma
Independent AI analyst, pk-sharma.com briefing

“Argues the real significance of Decision-1 is the architectural shift away from routing every task through general-purpose LLMs, not Nadella's personal endorsement, and cautions that Microsoft's internal benchmarks still need independent verification.”

Sophie Larsen
Author, remio.ai

“Positions Decision-1 as outperforming both general LLMs and rival decision models on latency and quality for structured decision tasks.”

Satya Nadella
CEO, Microsoft
The Crowd

“Introducing Microsoft-Decision-1, our new model for fast decision-making. It delivers top performance on structured decision tasks, outperforming both LLMs and other decision models in latency and quality. We're already testing it across Microsoft for everything from incident...”

@@satyanadella9556

“Microsoft-Decision-1 is live on OpenRouter. Microsoft's numbers: highest accuracy across 36 blind benchmarks (~150K questions), 4.5x faster than the runner-up, 35x faster than GPT-6 Sol, and decisions flip on just 1.3% of perturbed inputs. Post-trained from Qwen3.5-9B. $0.042/M”

@@OpenRouter524

“OpenAI Decision crushed by @Microsoft-Decision-1 at Space Invaders. Decision-1 decided 3.5x faster so it scored 3,320 and cleared 4 waves while OpenAI's decision API scored 1,300 taking over 300ms per decision on average. Run decision models ->”

@@atomic_chat_hq225

“Introducing Microsoft-Decision-1, our model for fast decision-making”

@u/Greedyanda527
Broadcast
Microsoft's New Decision Model: A Fresh AI Catalyst and the $537 Test | Oct 9

Microsoft's New Decision Model: A Fresh AI Catalyst and the $537 Test | Oct 9

Microsoft Decision-1: Same Price as Jev, Built on Qwen

Microsoft Decision-1: Same Price as Jev, Built on Qwen

Microsoft Just Beat GPT-6 With a Chinese Model

Microsoft Just Beat GPT-6 With a Chinese Model

Microsoft Decision-1 model launch — AI News | Agentic Brew