LLM cost planning autumn 2026: Sonnet 5 stays at $2/$10
Anthropic has cancelled the Claude Sonnet 5 price increase scheduled for 1 September 2026, making the introductory $2 input / $10 output per million tokens the permanent standard price. For teams running Claude on Azure through Microsoft Foundry, that removes a planned 50% jump from autumn budgets and reshapes the mid-tier price comparison against GPT-5.6 Terra and Gemini 3.1 Pro.
Azure OpenAI cost check: GPT-5.6 price cuts and PTU math
OpenAI cut GPT-5.6 Luna prices by 80 percent and Terra by 20 percent on 30 July 2026, and Microsoft confirmed the same decreases reach Azure OpenAI Global Standard deployments from 1 August. For Swedish and EU teams running these models on Azure, the cuts move the break-even point for PTU reservations, model routing and residency premiums, so the autumn budget math deserves a fresh pass before any new one-year commitments.
Claude Opus 5 for long-running agents: the cost math
Claude Opus 5 launched on 24 July 2026 at $5/$25 per million tokens with a 1M context window and day-one availability in Microsoft Foundry. For long-running agents the per-token price is the wrong unit: we work through cost per completed task against Sonnet 5 and GPT-5.6 Sol, and flag the EU data-residency caveat Swedish Azure teams need to check first.
Tokens are the new pricing lever: Gemini 3.6 Flash math
Google's 21 July release of Gemini 3.6 Flash pairs an output-price cut from $9.00 to $7.50 per million tokens with a claim of roughly 17% fewer output tokens per task, compounding to about 31% lower output cost for unchanged work. That combination makes per-million-token price sheets unreliable for model comparison, and Azure teams should measure cost per completed task instead.
Claude Sonnet 5 vs Opus 4.8 vs GPT-5.5: agent cost math
Anthropic launched Claude Sonnet 5 on 30 June 2026 at an introductory 2/10 dollars per million tokens, posting 63.2% on SWE-bench Pro and near-Opus agentic performance at 40-60% of the cost per task. We run the cost math against Opus 4.8 and GPT-5.5, set out a routing framework for when the cheap model wins, and draw the continuity lesson from the eighteen-day Fable and Mythos export-control pause.
Small Models in Production: When Phi-4 and 8B Llama Win
Frontier models are the default. Defaults are how teams overpay on LLM bills. Three workloads where small models (Phi-4, Llama 3.x 8B, Mistral Small) outperform on cost-per-decision without losing meaningfully on quality, three workloads where they do not, and the two-tier production pattern that cost-conscious teams converge on after a quarter of evaluation work.
Prompt Caching in 2026: Cut Azure OpenAI and Claude Costs
Prompt caching is the highest-ROI cost lever on long-context LLM workloads in 2026. Anthropic, OpenAI, and Azure OpenAI all offer it with different pricing and breakpoint semantics. A worked comparison of the three providers, the placement patterns that actually hit cache, where the cache silently goes cold, and a 30-minute audit that pays back.
AI Agent Cost Economics: Why 100x and How to Cut It
Agent loops cost 10x to 50x what a chatbot interaction costs; multi-agent systems add another order of magnitude. The cost compounding is structural, not a bug. The cost reduction is structural too. Decomposing where the tokens go and how to bring agent economics back from runaway to acceptable.
Cost-Optimizing Azure OpenAI: PTUs, Batch, Caching in 2026
A concrete playbook for reducing Azure OpenAI bills in 2026. Break-even math for Provisioned Throughput Units, prompt-cache economics, the Batch API 50 percent discount, Foundry IQ for retrieval, tiered model routing, and the telemetry that keeps the wins honest.
Microsoft Foundry: The AI Platform for the Agentic Era - Ignite 2025
From scientific research to enterprise AI transformation, discover how Microsoft Foundry unifies models from OpenAI, Anthropic, Cohere, Meta, and more into one secure platform. Learn intelligent model routing, cost optimization, and the game-changing Claude integration.
Fine-Tuning in Microsoft Foundry: Building Production-Ready AI Agents - Microsoft Ignite 2025
Microsoft Ignite BRK188: Fine-tuning in Microsoft Foundry transforms generic models into production-ready agents. Synthetic data generation, supervised + reinforcement fine-tuning, 40-90% cost reduction, 95%+ accuracy. Real-world results: 2M docs/day, $27M savings.
Running Open-Source AI Models at Scale: Azure Container Apps, AKS, and On-Premise Deployments - Microsoft Ignite 2025
Microsoft Ignite BRK117: Deploy open-source AI models (Llama 3.3, Mistral) with Azure Container Apps serverless GPUs, AKS with Kaido workflows, and on-premise infrastructure. Cost reduction 60-85%, data sovereignty, and hybrid architectures with Azure Arc.