# Claude Haiku 5.5 in Foundry: $0.10 and the 100K cliff

Claude Haiku 5.5 is GA in Microsoft Foundry at $0.10/$0.50 per million tokens, the same Global Standard list price as GPT-6 Luna, and 90% below Haiku 4.5 on paper. The bills diverge on three lines: a tier that charges 5x on every token once a prompt passes 100K, a tokenizer that counts the same text as 30% more tokens, and an EU Data Zone that Luna offers at $0.12/$0.60 while Claude on Foundry has none.

- Published: 2026-10-08 · Category: Azure & Cloud · Tags: Claude Haiku 5.5, Anthropic, Microsoft Foundry, Azure, GPT-6 Luna, LLM Cost, Prompt Caching, AI Agents, Model Migration, Data Residency
- Author: Technspire AB, Stockholm (https://technspire.com)
- Canonical: https://technspire.com/en/blog/claude-haiku-5-5-foundry-0-10-tier-100k-cliff-vs-gpt-6-luna

Claude Haiku 5.5 went generally available in Microsoft Foundry on 7 October 2026, the day Anthropic released it, at $0.10 per million input tokens and $0.50 per million output tokens. That is 90% below Claude Haiku 4.5 on both sides of the meter. It is also, to the cent, the Global Standard list price of GPT-6 Luna in the same catalog: $0.10 input, $0.01 cached input, $0.50 output. Two cheap-tier models, one price sheet. For an Azure team in Sweden the bills still diverge, for three reasons that sit below the headline: a price tier that quintuples at 100,000 prompt tokens, a tokenizer that counts the same text as roughly 30% more tokens, and an EU Data Zone that exists for Luna and not for Claude.

## What shipped on 7 October

The ID is `claude-haiku-5-5` on the Claude API, Google Cloud and Foundry, and `anthropic.claude-haiku-5-5` on Bedrock. It is a fixed ID with no date suffix. The context window grows from 200K to 1M tokens and max output from 64K to 128K. Adaptive thinking is on by default at `medium` effort, and this is the first Haiku with the full effort range from `low` to `max`. The reliable knowledge cutoff is June 2026, and Anthropic commits not to retire it before 7 October 2027.

Anthropic positions it for classification, routing, extraction and subagent work, and reports 72.4% on OSWorld 2.1, 39.2% on Terminal-Bench 4.0 and 45.9% on Humanity's Last Exam without tools. Asana is quoted at over 30% lower latency and up to 2.5x faster inference per agent turn; Box reports 11 points higher accuracy than Haiku 4.5 at about half the latency. Those are vendor-selected numbers. The one that matters is your own eval set at the effort level you intend to deploy.

In Foundry both hosting versions are GA on day one. Version 2, Hosted on Azure, runs inference end to end on Azure infrastructure and is what you get if you accept Default settings. Version 1, Hosted on Anthropic infrastructure, routes the request to Anthropic's own servers. Microsoft's own launch post puts the saving at "around 75% less than Claude Haiku 4.5 for most tasks" rather than Anthropic's 90%. The gap between those two numbers is the subject of this post.

GitHub Copilot got the model the same day for Pro, Pro+, Max, Business and Enterprise plans across VS Code, Visual Studio, JetBrains, Xcode, Eclipse, the CLI, the cloud agent and github.com, billed at provider list pricing under usage-based billing. New models are enabled automatically unless an administrator has turned off the global default, so check the model policy in Copilot settings this week if your organisation reviews models before developers use them.

### Price sheet, October 2026

| Model, per MTok | Input | 5m cache write | Cache read | Output |
| --- | --- | --- | --- | --- |
| Claude Haiku 5.5, prompt up to 100K | $0.10 | $0.125 | $0.01 | $0.50 |
| Claude Haiku 5.5, prompt over 100K | $0.50 | $0.625 | $0.05 | $2.50 |
| Claude Haiku 4.5 | $1.00 | $1.25 | $0.10 | $5.00 |
| GPT-6 Luna, Global Standard | $0.10 | none | $0.01 | $0.50 |
| GPT-6 Luna, EU Data Zone | $0.12 | none | $0.012 | $0.60 |
| Claude Sonnet 5.5 | $2.00 | $2.50 | $0.10 | $10.00 |

Haiku 5.5 is the only current Claude model priced by prompt length. Every other Claude 4.6-or-later model bills a 900K-token request at the same per-token rate as a 9K one. Above 100,000 prompt tokens, Haiku 5.5 charges five times the input rate and five times the output rate, on the whole request, not just the tokens past the line. Anthropic frames that tier as still 50% below Haiku 4.5, which is true per token. Luna's table has a long-context tier too, at $0.20 input and $0.75 output, but Microsoft's September post does not print the boundary; its Sol sibling switches at 272K input tokens, as we covered in [Sonnet 5.5 vs GPT-6.1 Sol on Azure](/en/blog/claude-sonnet-5-5-vs-gpt-6-1-sol-azure-cost-math). On Foundry all of these prices reach your invoice as Claude Consumption Units at $0.01 each, and the Message Batches discount of 50% is Claude API only. Luna's cache has no write charge in Microsoft's table, which matters more at this tier than it sounds.

## Cost math 1: the classification endpoint

Start with the workload Haiku exists for: 100,000 requests a month, each 20K input tokens and 1K output tokens as counted by Haiku 4.5, no cache reuse, thinking off. That is 2,000 million input tokens and 100 million output tokens on the old tokenizer. Haiku 5.5 uses the tokenizer introduced with Claude 4.7, so the same text counts as about 30% more tokens: 2,600 million in and 130 million out. The Luna rows assume equal token counts to Haiku 4.5, which is an assumption, because OpenAI's tokenizer is a third one.

| Monthly, 100K requests | Input | Output | Total |
| --- | --- | --- | --- |
| Haiku 4.5, 2,000M in / 100M out | $2,000 | $500 | $2,500 |
| Haiku 5.5, 2,600M in / 130M out, thinking disabled | $260 | $65 | $325 |
| Haiku 5.5, plus 1,000 thinking tokens per request | $260 | $115 | $375 |
| GPT-6 Luna, Global Standard | $200 | $50 | $250 |
| GPT-6 Luna, EU Data Zone | $240 | $60 | $300 |

The Haiku 4.5 to Haiku 5.5 saving is 87% once the tokenizer is priced in, not 90%. The 1,000-token thinking row is an illustrative assumption, not a measurement; it shows that even a modest amount of thinking at `low` effort only moves the saving to 85%. Unlike Opus 5.5, Haiku 5.5 does let you send `thinking: {"type": "disabled"}` at `high` effort or below, and Microsoft's Learn table confirms that value is accepted on Foundry. For a pure classifier, disable it and measure `usage.output_tokens` before and after.

The second thing the table shows is less comfortable. Luna with EU residency costs $300. Haiku 5.5 without any residency guarantee costs $325. At identical list prices the tokenizer difference is the whole margin, and it runs against Claude. If you chose Haiku for this tier on price alone, you now have two models at parity and only one of them processes in the EU.

## Cost math 2: the 100K cliff

Now a document workload: 100,000 requests a month, each a 150K-token prompt (a long contract, a support ticket history, a retrieved document set) and 2K output tokens. Haiku 4.5 would have accepted this inside its 200K window at $1 per million. Haiku 5.5 accepts it inside 1M, but every token of it is billed at the over-100K rate.

| Monthly, 100K requests | Input | Output | Total |
| --- | --- | --- | --- |
| Haiku 5.5, 150K in / 2K out | $7,500 | $500 | $8,000 |
| Haiku 5.5, same job trimmed to 95K in | $950 | $100 | $1,050 |
| GPT-6 Luna, Global Standard, 150K in | $1,500 | $100 | $1,600 |
| Claude Sonnet 5.5, 150K in | $30,000 | $2,000 | $32,000 |

Crossing the line costs 7.6 times more than staying under it, for a difference of 55K tokens. The Luna row assumes 150K is below Luna's unpublished long-context boundary, which holds if it matches Sol's 272K. Under that assumption Haiku 5.5 over the cliff costs five times Luna for the same request, and at parity under it. If your prompts cluster between 100K and 200K tokens, the engineering work that pays is not model selection. It is chunking, retrieval tightening or summarising the history so the prompt lands under 100K. On Sonnet 5.5 that work saves only the tokens you remove. On Haiku 5.5 it changes the rate on every token that remains.

Watch the tokenizer here too. A 90K-token prompt measured on Haiku 4.5 is roughly 117K tokens on Haiku 5.5. Workloads you believed were comfortably mid-sized can cross the line in the migration itself. Count with the token counting endpoint and `model` set to `claude-haiku-5-5`, not with the numbers in your Haiku 4.5 dashboards.

## Cost math 3: the cached agent session

Use the same 60-turn session we modelled for [Opus 5.5 in Foundry](/en/blog/claude-opus-5-5-foundry-opus-5-migration-cost-math): a cached prefix that averages 150K tokens per turn, giving 9M cache-read tokens, 3K new tokens written to the 5-minute cache per turn (180K in total) and 2K output tokens per turn including thinking (120K in total). Then run it again with the prefix held to 80K, giving 4.8M cache reads. Anthropic prices the tier by prompt length, and nothing on the pricing page exempts cached tokens from that count, so assume the 150K session pays the higher rate on every turn even though 95% of each prompt is a cache read.

| 60-turn session | Cache reads | Cache writes (180K) | Output (120K) | Total |
| --- | --- | --- | --- | --- |
| Haiku 5.5, 150K prefix (9M reads) | $0.45 | $0.11 | $0.30 | $0.86 |
| Haiku 5.5, 80K prefix (4.8M reads) | $0.05 | $0.02 | $0.06 | $0.13 |
| Haiku 4.5, 150K prefix | $0.90 | $0.23 | $0.60 | $1.73 |
| Sonnet 5.5, 150K prefix | $0.90 | $0.45 | $1.20 | $2.55 |

The long session is half the price of Haiku 4.5, which is exactly Anthropic's over-100K claim and nothing more. The short session is 13% of it. A subagent harness that compacts its context at 80K instead of letting it drift to 150K spends one seventh as much on Haiku 5.5, while on Sonnet 5.5 the same discipline would save under half. If Haiku 5.5 is going to be your parallel subagent model, which is the role Microsoft's post describes, then the orchestrator's compaction threshold is a pricing parameter. Set it below 100K and alert when a subagent crosses it.

## What breaks when you swap the deployment name

Anthropic's migration guide lists five breaking changes from Haiku 4.5 and three behaviour changes. On Foundry, one of the five does not apply.

- **Manual thinking budgets return 400.** `thinking: {"type": "enabled", "budget_tokens": N}` is rejected. Send `{"type": "adaptive"}` or omit the field, and steer depth with `output_config.effort`. Where Haiku 4.5 ran without thinking, pick `low` or disable it outright.
- **Sampling parameters return 400.** Any `temperature` other than 1, any `top_p` other than 0.99, any `top_k` at all, or both `temperature` and `top_p` together. Classification code that set `temperature: 0` for determinism breaks on the first request. Remove the parameters and constrain output with structured outputs or strict tools instead.
- **Assistant prefill returns 400.** Haiku 4.5 accepted a trailing assistant turn when thinking was off. Haiku 5.5 rejects it even with thinking disabled. Prefills used to force JSON move to `output_config.format`; prefills used as preambles move to the system prompt.
- **Thinking blocks are bound to the conversation and the account.** Replaying a thinking block after editing `system`, `tools` or an earlier message returns 400 on accounts created on or after 31 August 2026. Blocks replayed through a different account are silently dropped. Keep harnesses append-only.
- **The `computer_20250124` tool is rejected on the Claude API and Google Cloud only.** Foundry does not offer `computer_toolset_20260801` at all, and Anthropic's Foundry page says the beta computer use tool versions remain available there. A Foundry integration keeps working unchanged.

The behaviour changes catch more teams than the 400s. Responses can now begin with a `thinking` block, so code that reads `content[0].text` gets an empty string or a type error; select blocks by `type`. Thinking text is omitted by default, so set `thinking.display` to `"summarized"` if your logs relied on it. And thinking tokens count toward `max_tokens`, so a classifier with `max_tokens: 20` can stop after the thinking block with no answer. Raise the limit or disable thinking. Priority Tier is not supported on Haiku 5.5, which does not matter on Foundry, where it never existed, but matters if you were planning capacity across the Claude API too.

```
import AnthropicFoundry from "@anthropic-ai/foundry-sdk";
import { DefaultAzureCredential, getBearerTokenProvider } from "@azure/identity";

const client = new AnthropicFoundry({
  resource: "tsp-foundry-swc",
  azureADTokenProvider: getBearerTokenProvider(
    new DefaultAzureCredential(),
    "https://ai.azure.com/.default"
  ),
});

// Pure classifier: no thinking, no sampling params, schema-constrained output.
const res = await client.messages.create({
  model: "claude-haiku-5-5",                       // deployment name
  max_tokens: 64,
  thinking: { type: "disabled" },                  // allowed at effort high or below
  output_config: {
    effort: "low",
    format: { type: "json_schema", schema: TICKET_LABEL_SCHEMA },
  },
  system: STABLE_SYSTEM_PROMPT,                    // keep under 100K with the input
  messages,                                        // must end with a user turn
});

const text = res.content.find(b => b.type === "text")?.text;   // not res.content[0]
if (res.stop_reason === "refusal") {
  // no server-side fallback on Foundry: reroute to another deployment here
}
if (res.usage.input_tokens + (res.usage.cache_read_input_tokens ?? 0) > 100_000) {
  metrics.increment("haiku55.over_100k");          // you just paid 5x on this request
}
```

## The Swedish and EU angle

**Sweden Central is a deployment location, not a processing boundary.** Microsoft's region table lists Haiku 5.5 for Global Standard in Sweden Central, and its Claude models page lists Data Zone Standard for Haiku 5.5 in the US only. The Europe tab of the Data Zone table reads "Not available". A Global Standard deployment created in Sweden Central may process prompts in any Azure region. Luna, by contrast, is on Standard deployment across all 28 Global regions and both the US and EU Data Zones. For the cheap tier, the EU residency premium on Luna is 20% over Global. For Claude it is not for sale at any price.

**Check which version the portal offers you in Sweden Central.** As of 8 October, Learn's Global Standard region table lists Haiku 5.5 version 1 (Hosted on Anthropic) for swedencentral and shows version 2 (Hosted on Azure) only in the US Data Zone table, while Microsoft's catalog page says version 2 covers Global Standard too. That looks like documentation lag on launch day, but do not write "prompts remain within Azure" into a DPIA until the Model version dropdown in your own deployment shows version 2. The Anthropic-hosted version "might be processed outside Azure", in Microsoft's words, and Anthropic is the data processor under its own Data Processing Addendum in both versions.

**EU watermarking is on by default.** Microsoft's Claude page states that the Claude models in its table comply with the EU watermarking standard, with interwoven text watermarking applied server-side at generation time and no change to request or response shapes. That is a sharper position than the one we found for Azure OpenAI two days ago in [OpenAI's EU text watermark: Azure OpenAI is not covered yet](/en/blog/openai-eu-text-watermark-azure-openai-not-covered-yet). If your Article 50 transparency plan leans on provider-side marking, Haiku 5.5 on Foundry gives you something to cite today and Luna does not yet.

**No mandatory retention, and no server-side fallback.** Haiku 5.5 is not one of Anthropic's Covered Models, so the 30-day retention that applies to Fable 5.1 and Mythos 5.1 does not apply here. It does run safety classifiers that can return HTTP 200 with `stop_reason: "refusal"`, and Foundry has no server-side fallback, so your client reroutes. Cloud Solution Provider subscriptions remain unsupported for Claude in Foundry, which still catches Swedish mid-market companies buying Azure through a partner.

## Decision guide

- **Classification or extraction on Haiku 4.5 with prompts under 75K tokens?** Migrate. Expect roughly 85% lower cost after the tokenizer, not 90%. Disable thinking, strip sampling parameters, replace prefills with structured outputs.
- **Prompts between 75K and 200K tokens on the old tokenizer?** Recount first. Anything that lands over 100K on Haiku 5.5 pays 5x. Trim under the line, or price Luna for that workload.
- **Choosing between Haiku 5.5 and GPT-6 Luna at the same list price?** Under 100K tokens they are at parity, and Luna is cheaper per unit of text because of the tokenizer. Over 100K, Luna wins on price outright. Haiku 5.5 wins on effort control, 1M context, and the EU watermarking statement.
- **Need EU-only processing for the cheap tier?** Luna in the EU Data Zone at $0.12 and $0.60. Claude on Foundry cannot do it.
- **Using Haiku 5.5 as the subagent in an Opus 5.5 or Sonnet 5.5 harness?** Make the compaction threshold a pricing control. Below 100K the session costs $0.13; drift to 150K and it costs $0.86.
- **Buying Azure via CSP?** Arrange a supported subscription type before the pilot, as for every other Claude model in Foundry.

Whichever row you are in, count your prompts with the new tokenizer before you sign off the budget. The 90% headline is real for short prompts, 50% for long ones, and the line between them moved 30% closer the moment you changed the model ID.

## Sources

- [Anthropic: Introducing Claude Haiku 5.5 (pricing, benchmarks, customer quotes)](https://www.anthropic.com/claude-haiku-5-5)
- [Claude Platform Docs: Claude Haiku 5.5 overview (specs, long-context tiers, retirement date)](https://platform.claude.com/docs/en/models/haiku-5-5/overview)
- [Claude Platform Docs: What's new in Claude Haiku 5.5 (tokenizer, thinking, breaking changes)](https://platform.claude.com/docs/en/models/haiku-5-5/whats-new-haiku-5-5)
- [Claude Platform Docs: Claude Haiku 5.5 migration guide](https://platform.claude.com/docs/en/models/haiku-5-5/migration-guide)
- [Claude Platform Docs: Pricing (Haiku tiers, Foundry CCU billing, US Data Zone multiplier)](https://platform.claude.com/docs/en/about-claude/pricing)
- [Claude Platform Docs: Claude in Microsoft Foundry (hosting options, unsupported features)](https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry)
- [Claude Platform Docs: Data residency (inference_geo values, Foundry note)](https://platform.claude.com/docs/en/manage-claude/data-residency)
- [Microsoft Foundry Blog: Claude Haiku 5.5 is now available in Microsoft Foundry ("around 75% less")](https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/claude-haiku-5-5-is-now-available-in-microsoft-foundry/4562408)
- [Microsoft Learn: Claude models in Microsoft Foundry (hosting versions, thinking and effort tables, EU watermarking, Data Zone)](https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/claude-models)
- [Microsoft Learn: Foundry Models from partners and community (region tables, subscription types)](https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-from-partners)
- [Microsoft Foundry model catalog: claude-haiku-5-5](https://ai.azure.com/catalog/models/claude-haiku-5-5)
- [Microsoft Azure Blog: GPT-6 Astra, Sol and Luna in Microsoft Foundry (Luna pricing by deployment type)](https://azure.microsoft.com/en-us/blog/gpt-6-astra-sol-and-luna-for-production-agents-in-microsoft-foundry/)
- [GitHub Changelog: Claude Haiku 5.5 in GitHub Copilot](https://github.blog/changelog/2026-10-07-claude-haiku-5-5-in-github-copilot)

---

Technspire AB builds AI agents, Azure OpenAI solutions, and production web platforms for Swedish and EU enterprises. Book a call: https://calendly.com/technspire · hello@technspire.com · More articles: https://technspire.com/en/blog · Site overview for agents: https://technspire.com/llms.txt
