Back to all posts

cat posts/claude-sonnet-5-5-vs-gpt-6-1-sol-azure-cost-math.md --category "Azure & Cloud" --views 25

Sonnet 5.5 vs GPT-6.1 Sol on Azure: the $2/$10 cost math

Claude Sonnet 5.5 and GPT-6.1 Sol reached Microsoft Foundry a day apart at the same $2/$10 list price, yet a cache-heavy 60-turn agent session costs $3.45 on Sonnet 5.5 and $2.55 on GPT-6.1 Sol, or $3.06 in the EU Data Zone that Claude still lacks. Above 272K input tokens the order flips, because GPT-6.1 Sol doubles its input rate while Sonnet 5.5 keeps flat pricing to 1M tokens and can run with no up-front reasoning.

  • --author By Falak Mahmood
  • --date September 30, 2026
  • --read 14 min read
  • --views 25 views

Claude Sonnet 5.5 reached Microsoft Foundry on 28 September 2026 at $2 per million input tokens and $10 per million output tokens. Twenty-four hours later GPT-6.1 Sol went generally available in the same catalog at $2 and $10. Same list price, one day apart. For an Azure team in Sweden the choice now turns on four lines further down the price sheet: what a cache read costs, what happens above 272K input tokens, what EU processing adds, and whether the model can run without reasoning at all.

What shipped on 28 and 29 September

Claude Sonnet 5.5 uses the ID claude-sonnet-5-5 in Foundry. It keeps the 1M token context window at flat per-token pricing and 128K max output on the synchronous Messages API. The reliable knowledge cutoff is June 2026, and Anthropic commits not to retire the model before 28 September 2027. Prices are identical to Sonnet 5. The pitch is efficiency: Anthropic says the model generates output more than 30% faster than Sonnet 5 and needs fewer tokens per task. Microsoft's launch post lists it as GA in both hosting versions.

GPT-6.1 Sol is gpt-6.1-sol, version 2026-09-29, on Microsoft's list of Foundry Models sold by Azure. The context window is 1,050,000 tokens, split into 922,000 input and 128,000 output, with training data up to April 2026. Microsoft describes it as an upgrade to GPT-6 Sol with performance approaching GPT-6 Astra on agentic coding, computer use and professional work, and tells teams to use it as the new default for production agents. GPT-6 Sol went GA in Foundry on 22 September, so the model being replaced as the default is one week old.

There is no GPT-6.1 Astra to pair it with. TechCrunch, citing a Wall Street Journal report, writes that OpenAI scrapped that release after internal testing showed higher levels of deception and a tendency to proceed without asking the user for permission. We covered the related training pause in yesterday's agent sandbox audit.

Price sheet in Foundry, 30 September 2026

Model and deployment Input / MTok Cache write Cache read Output / MTok
Claude Sonnet 5.5, Global Standard$2.00$2.50$0.20$10.00
GPT-6.1 Sol, Global Standard, up to 272K$2.00$2.50$0.10$10.00
GPT-6.1 Sol, Global Standard, above 272K$4.00$5.00$0.20$15.00
GPT-6.1 Sol, EU Data Zone, up to 272K$2.40$3.00$0.12$12.00
GPT-6.1 Sol, EU Data Zone, above 272K$4.80$6.00$0.24$18.00
GPT-6 Sol, Global Standard, up to 272K$2.00$2.50$0.20$10.00

Microsoft's table labels the two GPT tiers short and long context without stating the boundary. OpenAI's model page does: prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output, for the full request. The Sonnet 5.5 cache write shown is the 5-minute one. A 1-hour write costs $4.00. Claude usage reaches your Azure invoice as Claude Consumption Units at a fixed $0.01 each.

Compare the second row with the last. The only price that moved between GPT-6 Sol and GPT-6.1 Sol is the cache read, halved from $0.20 to $0.10. That single cell decides the first comparison below.

Cost math 1: a cache-heavy agent session under 272K

We use the same 60-turn session as in the Claude Opus 5.5 migration math so the numbers line up across posts. The cached prefix averages 150K tokens per turn and never crosses 272K, which gives 9M cache-read tokens over the session. Each turn writes 3K new tokens to the cache (180K in total) and emits 2K output tokens including reasoning (120K in total). Initial writes and cache misses are ignored.

60-turn session Cache reads (9M) Cache writes (180K) Output (120K) Total
Claude Sonnet 5.5, Global$1.80$0.45$1.20$3.45
GPT-6 Sol, Global$1.80$0.45$1.20$3.45
GPT-6.1 Sol, Global$0.90$0.45$1.20$2.55
GPT-6.1 Sol, EU Data Zone$1.08$0.54$1.44$3.06

At identical token counts, GPT-6.1 Sol on Global Standard is 26% cheaper than Sonnet 5.5. The more interesting row for a Swedish buyer is the last one. GPT-6.1 Sol with EU-only processing, including the 20% Data Zone premium, still lands 11% below Sonnet 5.5 on Global Standard. A week ago, GPT-6 Sol against Sonnet 5 was a tie on Global Standard, and the same session with EU processing cost $4.14.

Cache reads dominate a long agent loop, which is why one halved cell moves the total this far. If your sessions are short or rarely reuse a prefix, the two models converge again.

Cost math 2: the same loop above 272K

Now load a large repository or a contract set so the prefix averages 400K tokens per turn. Every request is above OpenAI's threshold. Sixty turns give 24M cache-read tokens, with the same 180K written and 120K emitted.

60-turn session, 400K prefix Cache reads (24M) Cache writes (180K) Output (120K) Total
Claude Sonnet 5.5, Global$4.80$0.45$1.20$6.45
GPT-6.1 Sol, Global$4.80$0.90$1.80$7.50
GPT-6.1 Sol, EU Data Zone$5.76$1.08$2.16$9.00

The order flips. Sonnet 5.5 is 14% cheaper than GPT-6.1 Sol on Global Standard and 28% cheaper than the EU Data Zone deployment, because Anthropic bills a 900K-token request at the same per-token rate as a 9K-token one. The long-context tier also applies to the whole request, so a prompt of 280K input tokens pays double on all 280K. If a GPT-6.1 Sol agent hovers near the boundary, compacting context to stay under 272K is worth real money. On Sonnet 5.5 it saves only the tokens you remove.

Cost math 3: an endpoint that needs no reasoning

Take a high-volume extraction endpoint: 100,000 requests a month, 20K input tokens and 1K output tokens each, no cache reuse. That is 2,000 million input tokens and 100 million output tokens. On Global Standard both models cost $4,000 for input and $1,000 for output, $5,000 in total. In the EU Data Zone, GPT-6.1 Sol costs $6,000.

The difference sits in the reasoning floor. OpenAI's model page says GPT-6.1 Sol supports effort low through max, and that none and minimal are not supported. Reasoning tokens are billed as output. At this volume every 100 reasoning tokens per request adds $100 a month on Global Standard and $120 in the EU Data Zone. Sonnet 5.5 has a lower floor: thinking: {"type": "between_tools"} turns off up-front thinking, and on a request without tools the response contains only text.

Offline jobs add one more asymmetry. The Message Batches API is on Anthropic's list of features not supported for Claude in Foundry, so Sonnet 5.5 batch work pays the synchronous rate there. OpenAI lists batch pricing for GPT-6.1 Sol at 50% below standard, but Microsoft's launch post names only Standard and Provisioned Throughput deployments. Check the deployment type list in your own portal before you budget on a batch discount.

What the price sheet cannot tell you

All three tables assume both models consume the same number of tokens for the same work. They will not. The vendors use different tokenizers, so a million tokens is a different amount of text on each side, and Anthropic's pricing page notes that its current tokenizer produces about 30% more tokens for the same text than the one used through Sonnet 4.6. The two models also reason in different volumes at what sounds like the same effort level.

Anthropic's own launch page shows how wide the per-task spread is. Box reports that Sonnet 5.5 used 12% fewer total tokens than Sonnet 5. Balyasny Asset Management reports about 121K tokens per answer against 497K for Sonnet 5 on a private suite of 2,441 finance tasks, which is 76% fewer. Both are single-customer observations on their own workloads. Yours will sit somewhere else in that range, or outside it.

The published benchmarks do not settle it either. Anthropic's comparison table sets Sonnet 5.5 against GPT-6 Sol, the model OpenAI superseded the next day: 46.2% against 49.3% on FrontierCode 1.1, and 1844 against 1487 on GDPval-AA v2.1. No vendor has published Sonnet 5.5 against GPT-6.1 Sol. OpenAI's headline factuality figure, reported by TechCrunch, is that responses containing a factual error at low reasoning effort fall from 11.4% on GPT-6 Sol to 7.7%.

So run your own eval set on both deployments at the effort level you intend to ship, and record three numbers per task: whether it passed, total billed tokens by category from the usage object, and wall-clock time. Cost per completed task is the only figure worth putting in a procurement document.

What breaks on the way in

From Sonnet 5 to Sonnet 5.5 in Foundry

Anthropic lists five breaking changes. Two of them hit a typical Foundry integration, and one more changes the response shape without failing anything.

  • thinking: disabled returns 400. Send {"type": "between_tools"} instead. It is accepted at low, medium and high effort only, takes no display or budget_tokens field, and locks effort for the whole conversation. Microsoft Learn's thinking table still marks disabled and enabled as accepted for Sonnet 5.5. Build against Anthropic's migration guide.
  • Forced tool use returns 400. tool_choice of type any or tool is rejected, on the token counting endpoint too. Keep auto, mark the tool strict: true (at most 20 strict tools per request), and say in the prompt when the tool must be called.
  • Notes between tool calls arrive as thinking blocks. With adaptive thinking their text is empty at the default display: "omitted", so an agent UI goes quiet between steps with no error. Set display to "summarized", or use between_tools, where the text comes back.

The other three matter less here. Thinking blocks are now signed over the conversation before them, but Anthropic names the Claude API, Amazon Bedrock and Google Cloud for default enforcement and does not name Foundry. Keep your harness append-only anyway. The computer_20251124 rejection applies to the Claude API and Google Cloud, and the advisor tool is not available in Foundry at all.

Two quieter changes deserve a line in the migration ticket. Effort levels are recalibrated, so a setting carried over from Sonnet 5 does not buy the same amount of thinking, and Anthropic recommends re-running the effort sweep. Sonnet 5.5 also declines in more categories than Sonnet 5. A refusal returns HTTP 200 with stop_reason: "refusal", and server-side fallback is not supported in Foundry, so your client has to reroute.

import AnthropicFoundry from "@anthropic-ai/foundry-sdk";

const client = new AnthropicFoundry({
  resource: "tsp-foundry-swc",
  azureADTokenProvider: tokenProvider,
});

// Sonnet 5 sent thinking: { type: "disabled" }. On Sonnet 5.5 that is a 400.
const res = await client.messages.create({
  model: "claude-sonnet-5-5",             // deployment name
  max_tokens: 2000,
  thinking: { type: "between_tools" },    // lowest setting, no other fields allowed
  output_config: { effort: "low" },       // low, medium or high; fixed per conversation
  system: EXTRACTION_PROMPT,
  messages,
});

const text = res.content.filter(b => b.type === "text");   // select by type, not content[0]
if (res.stop_reason === "refusal") {
  // no server-side fallback in Foundry: retry on another deployment
  console.warn(res.stop_details?.category);
}

From GPT-6 Sol to GPT-6.1 Sol on Azure

We found no Microsoft migration guide for a model that replaces a one-week-old predecessor, but the capability table on Microsoft Learn carries one difference worth testing. For gpt-6-sol the row reads functions, tools and parallel tool calling. For gpt-6.1-sol the same row adds "Responses API only". If your agent calls tools through Chat Completions, run the integration tests before you change the deployment. Teams still on GPT-5.6 Sol should start from our GPT-5.6 Sol migration checklist.

Foundry differences that outlast the price

In Foundry Claude Sonnet 5.5 GPT-6.1 Sol
SellerAnthropic via Azure Marketplace, billed in CCUsAzure, as a Foundry Model sold by Azure
EU Data ZoneNot availableStandard, at a 20% premium
Reserved capacityNone listedProvisioned Throughput in Global and US Data Zone at launch
Context pricingFlat to 1M tokens2x input, 1.5x output above 272K input tokens
Lowest reasoning settingbetween_tools, no up-front thinkingEffort low
Knowledge cutoffJune 2026April 2026
Content filteringNot built in at deployment timeFoundry content filters and guardrails
Batch discountMessage Batches API not supportedListed by OpenAI, not named in Microsoft's launch post

The content filtering row is easy to miss. Microsoft Learn states that Foundry does not provide built-in content filtering for Claude models at deployment time and tells you to configure AI content safety during inference. If your governance baseline assumes every Foundry deployment sits behind the platform filter, a Claude deployment needs its own control.

The Swedish and EU angle

Only one of the two can be pinned to the EU. GPT-6.1 Sol is available as Data Zone Standard (EU) from day one at $2.40 and $12.00. Claude's Data Zone Standard table on Microsoft Learn lists US regions, and the Europe tab reads "Not available". A Sonnet 5.5 deployment created in Sweden Central is a Global Standard deployment, and prompts may be processed in any Azure region that hosts the model. We covered how the EU premium went from 10% to 20% in our Swedish cost math on the EU Data Zone. After this week, paying that premium on GPT-6.1 Sol can still be cheaper than not paying it on Sonnet 5.5.

Three official pages give three answers on where Sonnet 5.5 runs. Microsoft's launch post lists a Hosted on Azure version for Global Standard and US Data Zone. Anthropic's Foundry page says Sonnet 5.5 supports only Global Standard. Microsoft Learn's region table, as of today, lists only version 1 under Global Standard in Sweden Central, and version 1 is the one hosted on Anthropic's infrastructure outside Azure. This matters for a DPIA because Anthropic's statement that prompts and completions remain within Azure applies to the Azure-hosted version only. Open the deployment in the portal and read the model version before you write the data flow down. Version 2 is Hosted on Azure.

The contract differs. GPT-6.1 Sol is bought as a first-party Azure service. Claude is a Non-Microsoft Product under the Product Terms, and Anthropic's documentation says customers using Claude through Foundry are subject to Anthropic's data use terms. For a Swedish public-sector buyer that is a second supplier to assess.

CSP subscriptions get GPT-6.1 Sol only. Microsoft lists Cloud Solution Provider subscriptions as unsupported for Claude and advises CSP customers to use models offered as a first-party consumption service. Many Swedish mid-market companies buy Azure through a CSP partner.

Reserved capacity with EU processing is missing for both. Provisioned Throughput for GPT-6.1 Sol launched in Global and US Data Zone deployments. Microsoft's 22 September announcement did include the EU Data Zone for GPT-6 Sol Provisioned Throughput, so a latency-critical EU workload may have a reason to stay on the older model for now.

Decision guide

  • Need EU-only processing? GPT-6.1 Sol in the EU Data Zone. Sonnet 5.5 cannot offer it in Foundry at any price.
  • Cache-heavy agent loops that stay under 272K? GPT-6.1 Sol is cheaper on the price sheet: 26% on Global Standard, 11% with EU processing. Confirm with tokens per completed task.
  • Context regularly above 272K? Sonnet 5.5, or compact the GPT-6.1 Sol context below the threshold. Flat pricing is worth 14 to 28% on our session model.
  • High-volume endpoint with no reasoning need? Sonnet 5.5 with between_tools. If EU processing is mandatory, measure GPT-6.1 Sol at low and price the reasoning tokens.
  • Already on Sonnet 5? The per-token price is unchanged, so the saving is whatever token reduction you measure. Fix thinking and tool_choice first.
  • Already on GPT-6 Sol? The swap halves your cache read rate. Test tool calling on Chat Completions before cutover.
  • Buying Azure through a CSP? GPT-6.1 Sol, unless you arrange a separate supported subscription for Claude.

Whichever model you pick as primary, deploy the other one as a tested fallback. Same-price models from two vendors make a second deployment cheap to keep warm, and the pattern is in our multi-provider failover plan.

subscribe # the AI news that matters, minus the noise

Book a Call

Sources

Tags

Related posts