# Microsoft-Decision-1 in Foundry: $0.042 and the EU zone

Microsoft-Decision-1 went GA in Microsoft Foundry on 9 October 2026 at $0.042 per million input tokens with free output, returning calibrated probabilities over fixed options instead of text, and it launched with an EU Data Zone at the same price as the US zone. We compare it with OpenAI's Decisions API and H2O-Lightning-4B, run the cost math for ticket routing and agent step-gating, and set out what Swedish buyers must record about its Qwen3.5-9B base.

- Published: 2026-10-10 · Category: Azure & Cloud · Tags: Microsoft-Decision-1, Microsoft Foundry, Decision Models, Azure, AI Agents, Classification, EU Data Zone, Data Residency, OpenAI Decisions API, LLM Cost
- Author: Technspire AB, Stockholm (https://technspire.com)
- Canonical: https://technspire.com/en/blog/microsoft-decision-1-foundry-0-042-eu-data-zone-decision-models

Microsoft-Decision-1 went generally available in Microsoft Foundry on 9 October 2026 at $0.042 per million input tokens, with no charge for output. It does not generate text. You hand it a situation and a question with a fixed set of answers, and it returns a calibrated probability for each answer in a single forward pass. Microsoft reports a median latency 35 times lower than GPT-6 Sol on the same decisions, and 2.5 times lower than H2O-Lightning-4B, the open-weight model that defined this category. For a Swedish team on Azure the detail that matters most sits in the pricing table: the model launched with an EU Data Zone deployment on day one, at the same price as the US zone. That is not how new Foundry models usually arrive.

## What a decision model is, and what it is not

Every production agent spends most of its calls on choices rather than prose. Which team gets this ticket. Is this tool call safe to run. Does this draft meet the rubric. Did the retrieval return anything relevant. Teams have been answering those questions by prompting a general LLM for a one-word JSON answer, which works, but pays for a model built to write paragraphs and waits for it to finish thinking. A decision model inverts that. Microsoft's own description: "Unlike LLMs, which are designed to generate text or reason through complex problems, decision models are purpose-built to deliver structured outputs that software can immediately act on."

The output is a probability distribution over the options you supplied. Microsoft-Decision-1 supports yes/no questions, multiple choice, ratings on a scale, and rubric-based grading. Calibration is the point: as the launch post puts it, "a 90% prediction should be right about nine times out of 10." That property is what lets you set a confidence threshold and route everything below it to a human, something a free-text "billing" from a chat model cannot give you without a second call.

The Foundry model card is equally clear about what it is not for: text generation, open-ended question answering, conversation, translation and summarisation are all out of scope. So are "consequential decisions about people" in domains such as credit, employment or healthcare when the model is the sole decision-maker. Keep that sentence; it reappears in the EU section.

## What Microsoft shipped on 9 October

- **Base and training.** The model is Alibaba's open-weight Qwen3.5-9B, post-trained by Microsoft on public datasets that passed Microsoft's Open Data process plus synthetic data. Microsoft says it "will soon rebase it on other models, including Microsoft AI (MAI) and OpenAI." The weights are not published.
- **Limits.** Context window of 32,768 tokens. Text input only; output is JSON. No image input, which both OpenAI's and H2O's equivalents offer.
- **Deployment.** A Direct from Azure model, Microsoft-managed, GA at version 1. The Tech Community post lists two deployment types, US Data Zone and EU Data Zone, both at $0.042 per million input tokens and no output charge. No Global Standard row and no batch tier appear in the launch material.
- **Benchmarks.** Microsoft reports the highest accuracy in its own 36-benchmark comparison of nearly 150,000 questions kept blind from training. The individual benchmark names and per-model scores are not in either launch post, so treat the accuracy claim as vendor-reported until a third-party leaderboard runs it. The latency claims are more concrete: P50 roughly 35 times faster than GPT-6 Sol and 2.5 times faster than H2O-Lightning-4B v1.1.
- **Consistency.** Across eight request perturbations the model changed its decision on 1.3% of inputs on average, with zero flips when option descriptions were paraphrased or options reversed or shuffled. Safety testing covered 5,250 requests across 11 benchmarks including jailbreak and prompt-injection sets.
- **Internal use.** Microsoft cites Xbox Research categorising more than 10,000 pieces of player feedback at "quality competitive with GPT-5 while running 80–100 times faster", plus use in Copilot response grading and adaptive replanning in Microsoft Discovery.
- **Availability elsewhere.** The launch post names OpenRouter as a second channel. The OpenRouter listing returned a 404 when we checked on 10 October, and the Tech Community version says "coming soon", so plan on Foundry for now.

### The API shape

This is not a chat completion and your OpenAI-compatible SDK will not talk to it. The endpoint is a provider route on your Foundry resource, authenticated with the usual Entra ID token for the Cognitive Services scope. The request carries the text to judge in `state` and one or more named questions, each with a type, instructions and a map of option keys to descriptions. Microsoft's own example routes a support ticket:

```
POST {AZURE_ENDPOINT}/providers/microsoft/v1/systemone
Authorization: Bearer <Entra ID token, scope https://cognitiveservices.azure.com/.default>

{
  "model": "<deployment name>",
  "state": "I cannot sign in after resetting my password.",
  "questions": {
    "team": {
      "type": "choice",
      "instructions": "Which team should handle this customer support request?",
      "criteria": {
        "billing":   "Charges, invoices, refunds, or subscription payments",
        "technical": "Software errors, bugs, or integration failures",
        "account":   "Sign-in, password, or account-access problems"
      }
    }
  }
}

// response (abridged)
{ "answers": { "team": { "type": "choice", "choice": "account", ... } } }
```

The option descriptions are part of the prompt and are billed as input tokens, so a question with a dozen long criteria costs more per call than the ticket itself. That is still cheap at $0.042, but it changes how you write option text: short, distinct descriptions beat paragraphs, and the perturbation results suggest paraphrasing them will not move the answer much.

## The category it joins: four decision models in one week

Decision-1 is not a lone experiment. In the eight days before it shipped, the category went from a research niche to a product line at three vendors.

| Model | Base | Input price per MTok | Images | Where it runs | EU processing |
| --- | --- | --- | --- | --- | --- |
| Microsoft-Decision-1 (GA 9 Oct) | Qwen3.5-9B, post-trained | $0.042, output free | No | Microsoft Foundry | EU Data Zone, same price |
| OpenAI Decisions API (beta 6 Oct) | gpt-6-luna | $0.10, no output or cache charges | Yes, inline base64 | OpenAI API only | Europe residency (EEA + Switzerland), ZDR |
| H2O-Lightning-4B v1.1 (7 Oct) | Qwen3.5-4B | Apache 2.0 weights, your GPU | Yes, up to 4 | Self-hosted | Wherever you run it |
| GPT-6 Luna via chat (for reference) | gpt-6-luna | $0.10 Global, $0.12 EU zone, plus output | Yes | Foundry, OpenAI | EU Data Zone at +20% |

Three observations from that table. First, OpenAI's Decisions API is the direct competitor and it is not on Azure. It runs GPT-6 Luna through a dedicated `/v1/decisions` endpoint with predicate, choice and score question types, bills input only at $0.10, and claims answers about ten times faster than the Responses API. It is in public beta with GA "in the coming weeks". If your organisation is Azure-only by policy, that option does not exist for you yet, and there is no announcement of it coming to Foundry. Second, H2O-Lightning-4B is the open-weight route: Apache 2.0, a 4B base, 29 ms median latency on an H100 by H2O's own measurement, and top of the JevBench composite as of 7 October. It is the choice when the data cannot leave your own hardware. Third, two of the three hosted or open options are built on Qwen 3.5, Alibaba's open-weight family. That is a provenance fact some Swedish buyers will need to write down, and we come back to it below.

## Cost math: two workloads

Prices this low make the arithmetic look trivial, so run it at two scales. Prices are list, Foundry EU Data Zone where one exists, and exclude the OpenAI Decisions API from the Azure rows because it is not in Foundry.

### Workload 1: support ticket routing

500,000 tickets a month, each call about 600 input tokens once the ticket, instructions and option descriptions are counted. That is 300 million input tokens.

- **Microsoft-Decision-1, EU Data Zone:** 300M × $0.042 = about $12.60 a month.
- **GPT-6 Luna chat, EU Data Zone, reasoning minimal, 10-token JSON answer:** input $36 plus output under $3, about $39. Turn reasoning on and the output side grows by an order of magnitude.
- **Claude Haiku 5.5, Global (no EU zone exists):** input about $39 once the newer tokenizer's roughly 30% token inflation is counted, plus a few dollars of output. We covered that tokenizer effect in [Claude Haiku 5.5 in Foundry: $0.10 and the 100K cliff](/en/blog/claude-haiku-5-5-foundry-0-10-tier-100k-cliff-vs-gpt-6-luna).

At this scale nobody is choosing on price. A three-fold gap on a $40 line item is invisible on an Azure invoice. What you are buying is the latency, the calibrated probability and the EU zone, in that order.

### Workload 2: step-gating an agent fleet

Now put the model where Microsoft's use-case list points: agent controls, verification, model routing. An agent platform that checks every proposed tool call against a policy rubric might make 20 million checks a month, each carrying about 2,000 tokens of tool call, context and rubric. That is 40 billion input tokens.

| Option | Input cost | Output cost | Monthly |
| --- | --- | --- | --- |
| Microsoft-Decision-1, EU Data Zone | $1,680 | $0 | $1,680 |
| GPT-6 Luna chat, EU Data Zone, 20 output tokens | $4,800 | $240 | $5,040 |
| Claude Haiku 5.5, Global, 20 output tokens | $5,200 | $260 | $5,460 |
| OpenAI Decisions API (not on Azure) | $4,000 | $0 | $4,000 |

Here the gap is $3,000 to $4,000 a month, and it compounds with the latency figure. A check that sits inside every agent step adds its P50 to every step. If Microsoft's 35x holds on your traffic, a gate that cost an agent two seconds per action costs it a fraction of that, which changes what you can afford to gate at all. That is the real argument for the category: policy checks you previously sampled become checks you run on every call.

## Decision guide: which model for which decision

- **Pick Microsoft-Decision-1** when the option set is fixed and small, the volume is high, you want a probability to threshold on, the input is text under 32K tokens, and the data must be processed inside the EU boundary on Azure. Routing, triage, relevance scoring, rubric grading of other models' output and tool-call policy gates all fit.
- **Pick the OpenAI Decisions API** when you need image input, you are not bound to Azure, and the judgement is hard enough to want a reasoning-grade model behind it. Accept beta status and the 2.4x input price.
- **Pick H2O-Lightning-4B** when the data cannot leave your own hardware, you need images, or you want to fine-tune. Budget the GPU and the ops.
- **Stay with a general LLM** when the answer is not enumerable in advance, when you need the explanation as well as the verdict, or when the task is extraction rather than choice. Our [Azure routing playbook](/en/blog/gpt-5-6-sol-vs-terra-vs-luna-a) still applies for the generation side; a decision model is what sits in front of it choosing the lane.
- **Do not use any of them as a sandbox.** A classifier that says a tool call looks safe is a signal, not a boundary. We made that argument in [A classifier is not a sandbox](/en/blog/coding-agent-sandboxing-auto-mode-break), and a faster classifier does not change it.

## Five things to check before you ship it

- **1. Calibrate on your own labels.** Microsoft's calibration claim is measured on its blind benchmark set. Take a thousand historically labelled tickets or tool calls, run them through, and plot predicted probability against observed accuracy in ten buckets. Set your human-review threshold from that plot, not from the launch post.
- **2. Pin the version and plan for the rebase.** Microsoft has said it will move the model from Qwen3.5-9B to MAI or OpenAI bases. When that lands, the calibration curve you measured in step 1 is void. Pin version 1 in your deployment and re-run the calibration set before promoting any new version.
- **3. Count the option tokens.** Instructions and criteria descriptions are billed input on every call. Put the stable part of your rubric under test: shorter descriptions that keep the 1.3% flip rate are free money.
- **4. Check the deployment type you actually got.** The launch lists Data Zone deployments only. Confirm in the portal that your deployment SKU is DataZoneStandard in an EU region, and apply the Azure Policy that denies other SKUs on the resource if your residency commitment requires it.
- **5. Keep a fallback lane.** Low-confidence answers need somewhere to go. Route them to a human queue or to a general model with reasoning on, and log the probability alongside the final decision so you can audit the threshold later.

## The Swedish and EU angle

**The EU Data Zone at launch is the headline for Swedish buyers.** Microsoft's documented order for new models is Global first, Data Zone later, geography-based last, and the EU zone normally carries a premium: GPT-6 Luna costs $0.12 in the EU zone against $0.10 Global. Decision-1 skipped the Global stage and priced both zones identically. For a Data Zone deployment, Microsoft processes prompts and responses only within the Azure EU Data Boundary, which can include EFTA countries such as Norway and Switzerland alongside the member states, and data at rest stays in your chosen geography. For a Swedish organisation that has refused Claude Haiku 5.5 because it has no EU zone, this is the first cheap-tier model in Foundry that clears the residency bar on day one.

**Provenance needs a sentence in your model register.** The model is Microsoft-hosted, Microsoft-post-trained and sold under Microsoft's terms, and no request data goes to Alibaba. The base weights are still Qwen3.5-9B under Apache 2.0. Several Swedish public-sector and defence-adjacent buyers now ask vendors to document the origin of model weights, and "Microsoft model" is not a complete answer to that question. Write down the base, the licence and Microsoft's stated plan to rebase, and you have answered it properly.

**The AI Act applies to the decision, not the model size.** A model that routes support tickets is minimal-risk. The same model deciding who gets a loan interview or which job applicant advances is an Annex III high-risk use, and Microsoft's own out-of-scope note excludes exactly those cases when the model acts alone. Under the Digital Omnibus that entered into force on 27 July 2026, the stand-alone Annex III obligations now apply from 2 December 2027, which gives you time but not a pass. GDPR Article 22 on solely automated decisions with legal or similarly significant effects has no deferral at all. The calibrated probability is your friend here: a documented threshold below which a human decides is the simplest honest implementation of meaningful human involvement, and it is a design you can show an auditor.

**Procurement gets a real alternative to building.** Many Swedish teams have run their own fine-tuned BERT-class classifiers for years precisely because LLM calls were too slow and too expensive for routing. A hosted decision model at this price, inside the EU boundary, with a probability output, is the first managed service that competes with that pattern on its own terms. The trade is control for operations: you lose the ability to retrain on your data, and you gain never patching a GPU node again.

## Conclusion

Microsoft-Decision-1 is a narrow model with a precise job, and Microsoft priced it like one. Treat the accuracy claim as unverified until independent numbers arrive, and treat the latency, the calibrated output and the EU Data Zone as the reasons to run a pilot this month. Calibrate on your own data, pin version 1, keep a fallback lane for low confidence, and put the model in front of every agent action you previously only sampled. If the Qwen base is a procurement problem, say so now, because Microsoft has already told you the base will change.

## Sources

- [Microsoft: Introducing Microsoft-Decision-1, our model for fast decision-making (9 October 2026)](https://commandline.microsoft.com/microsoft-decision-1-model-foundry/)
- [Microsoft Tech Community: Introducing Microsoft-Decision-1 in Microsoft Foundry for decision and classification workloads](https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/introducing-microsoft-decision-1-in-microsoft-foundry-for-decision-and-classific/4562742)
- [Microsoft Foundry model catalog: Microsoft-Decision-1 model card](https://ai.azure.com/catalog/models/Microsoft-Decision-1)
- [Microsoft Learn: Understanding deployment types in Microsoft Foundry Models](https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/deployment-types)
- [OpenAI: Decisions API guide](https://developers.openai.com/api/docs/guides/decisions)
- [OpenAI API changelog, 6 October 2026 entry](https://developers.openai.com/api/docs/changelog)
- [Hugging Face: h2oai/h2o-lightning-4b model card](https://huggingface.co/h2oai/h2o-lightning-4b)
- [Hugging Face: Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B)

---

Technspire AB builds AI agents, Azure OpenAI solutions, and production web platforms for Swedish and EU enterprises. Book a call: https://calendly.com/technspire · hello@technspire.com · More articles: https://technspire.com/en/blog · Site overview for agents: https://technspire.com/llms.txt
