2026-10-11 18:31 UTC
DANGMUAAI & Developer Tools, Decoded
BackAI Models

Microsoft's Decision-1: $0.042 per Million Input Tokens

Microsoft shipped a decision-scoring model on October 9 at $0.042 per million input tokens, output free. Every benchmark figure is Microsoft's own.

DangMua EditorialOct 11, 20266 min read
Microsoft's Decision-1: $0.042 per Million Input Tokens

Microsoft released Microsoft-Decision-1 on October 9, a decision-scoring model priced at $0.042 per million input tokens with output tokens free.

That price is the news. If a model can decide which tool to call, how risky an operation is, or whether an agent's output passes a rubric, and it costs almost nothing to run, the economics of routing through a general-purpose LLM stop making sense for a large class of steps inside agent pipelines.

What the model actually does

Decision-1 does not generate text. It takes a state description plus a fixed set of answer options and returns a calibrated probability for each option, as JSON. The current version handles yes/no questions, multiple-choice selections, ratings, and rubric-based grading of AI responses or agent actions.

Microsoft post-trained Alibaba's open-weight Qwen3.5-9B for single-pass decision scoring. The company says it plans to rebase the model on other foundations later, including Microsoft AI (MAI) and OpenAI models. Context window is 32,768 tokens, text only. The weights are not released — it ships as a hosted API through Microsoft Foundry and OpenRouter, in public preview. Achint Srivastava, VP of Software Engineering in the Office of the CTO, announced it.

The design goals Microsoft named are worth reading as a spec: low latency, generalization across held-out benchmarks, robustness to paraphrasing and reordering of options, well-calibrated probabilities, and safety filtering that refuses harmful requests.

The numbers, and who produced them

Microsoft reports an internal evaluation across 36 benchmarks covering nearly 150,000 questions held blind from training. Decision-1 took the highest average accuracy at 83.5%, ahead of Jev 1.13.0 at 82.3%, Quyet-1.0-Large at 81.9% and GPT-6 Luna Decisions at 79.4%.

Speed is the wider gap. Microsoft puts P50 latency at roughly 85 ms in its tests, calls the model about 2.5 times faster than the next-best decision model and roughly 35 times faster than GPT-6 Sol on the same tasks. Calibration scored 92.2 against a perfect 100. On robustness, decisions changed on 1.3% of perturbations on average, with zero flips when option descriptions were paraphrased or options reordered.

Every one of those figures comes from Microsoft's own reporting. No independent third-party benchmark was available at announcement. The accuracy spread at the top — 83.5% against 82.3% — is 1.2 points between a vendor's model and its closest named competitor, measured by the vendor. Treat the latency and price claims as the load-bearing ones and the accuracy ranking as provisional.

Where Microsoft is already running it

The internal deployments Microsoft disclosed are more useful than the benchmark table, because they name the workload:

  • Xbox Research classified more than 10,000 pieces of open-ended feedback into researcher-defined themes, reporting results competitive with GPT-6 Sol while running more than 14 times faster and 200 times less expensive.
  • Copilot used it for quality measurement of chat and agent responses, which Microsoft describes as competitive with GPT-5.6 Luna at roughly 100 times the speed.
  • Microsoft Discovery used it for adaptive replanning in scientific experiments, reporting greater consistency and nearly four times the overall speed versus an LLM-based approach.
  • On-call engineers tested it for retrieving relevant knowledge during live incidents.

Three of the four are evaluation and classification jobs, not live user-facing routing. That is a reasonable place to start: grading and theme-classification are tasks where a wrong answer costs you a mislabeled row, not a failed transaction.

Decision-1 versus asking an LLM

The category Decision-1 is entering already has a working pattern in open source. A recent LangGraph.js walkthrough built an agent around Jev, which its author describes as a "System One decision model" that takes a state and typed questions and returns structured answers with calibrated probabilities, without generating any text. Jev exposes three question types — Choice, Score and Noul — selecting from options, scoring an ordered scale, and producing a yes/no probability.

That walkthrough shows what the structured output buys you in code. Instead of parsing a sentence like "I think we should probably execute the request", the application branches on a number:

if (safety.noul >= 0.9) { return "execute"; }
if (safety.noul >= 0.6) { return "review"; }
return "blocked";

That snippet is from the Jev sample, not from Decision-1's API, but the shape generalizes to any scorer returning calibrated probabilities. Two design rules in that walkthrough apply regardless of which vendor you pick. First, the scorer provides judgment and does not execute anything — the application decides what the judgment means operationally. Second, do not push everything into the model: the sample keeps a deterministic prefilter in ordinary code for null checks and length limits, because, as its author puts it, "there are many things that ordinary code does better."

What to check before wiring it in

Three constraints decide whether this is usable for you.

Text only. No multimodal input. If your routing decisions depend on screenshots or document images, Decision-1 is out and some competing decision APIs that accept images are not.

Closed weights, hosted only. There is no self-hosting path. For teams already in Azure this is a feature; for anyone with data-residency constraints that rule out a hosted scorer, it is disqualifying in a way the open-weight alternatives are not.

Your decision distribution is not Microsoft's. The 36-benchmark result says the model generalizes across the tasks Microsoft chose. It does not say it generalizes to your intent taxonomy or your risk rubric. Build a held-out set from your own production traffic and measure against it before the model gates anything that touches money or user data.

The thing to watch is whether a third party reproduces the latency claim. An 85 ms P50 at $0.042 per million input tokens with free output is cheap enough to insert into every step of a multi-stage pipeline without moving the cost or latency budget — which is exactly the claim that makes this category interesting, and exactly the one no one outside Microsoft has checked yet.

Sources

  • "Microsoft Launches Decision-1: Fast Decision-Scoring Model Built on Qwen3.5-9B" — TechPulse, Dev.to
  • "Building an AI Agent Decision Layer with Jev and LangGraph" — Peter Saktor, Dev.to

More from DangMua