DeepSeek's 4x price rise: rethinking cheap open models
DeepSeek's V4-Pro reached general availability on 13 August 2026 with a 1M-token context and strong agent benchmarks, and three days later its peak-hour output price rose from a flat $0.87 to $3.96 per million tokens. For EU teams that built agent cost models around ultra-cheap open-weight APIs, the arithmetic, the data governance questions and the Azure hosting options all deserve a fresh look.
Workhorse shootout: Gemini 3.7 Flash, GPT-5.6 Luna, Sonnet 5
Google shipped gemini-3.7-flash as generally available on 13 August 2026 at an introductory $0.75/$3.75 per million tokens, two weeks after OpenAI cut GPT-5.6 Luna by 80 percent and days after Anthropic locked Claude Sonnet 5 at $2/$10 permanently. We compare the three workhorse models on list price, context, Azure availability and EU residency, and show why cost per completed task beats cost per token.
LLM cost planning autumn 2026: Sonnet 5 stays at $2/$10
Anthropic has cancelled the Claude Sonnet 5 price increase scheduled for 1 September 2026, making the introductory $2 input / $10 output per million tokens the permanent standard price. For teams running Claude on Azure through Microsoft Foundry, that removes a planned 50% jump from autumn budgets and reshapes the mid-tier price comparison against GPT-5.6 Terra and Gemini 3.1 Pro.
Claude Opus 5 for long-running agents: the cost math
Claude Opus 5 launched on 24 July 2026 at $5/$25 per million tokens with a 1M context window and day-one availability in Microsoft Foundry. For long-running agents the per-token price is the wrong unit: we work through cost per completed task against Sonnet 5 and GPT-5.6 Sol, and flag the EU data-residency caveat Swedish Azure teams need to check first.
Tokens are the new pricing lever: Gemini 3.6 Flash math
Google's 21 July release of Gemini 3.6 Flash pairs an output-price cut from $9.00 to $7.50 per million tokens with a claim of roughly 17% fewer output tokens per task, compounding to about 31% lower output cost for unchanged work. That combination makes per-million-token price sheets unreliable for model comparison, and Azure teams should measure cost per completed task instead.
Running GPT-5.6 the enterprise way on Microsoft Foundry
GPT-5.6 (Sol, Terra, Luna) went GA in Microsoft Foundry on 9 July 2026, day-and-date with OpenAI, alongside a new Asia-Pacific Data Zone and a hosted agents runtime with VNet integration. A practical guide for Swedish and EU Azure teams: choosing between the three models, picking Global Standard versus EU Data Zone versus PTUs, worked cost math on the launch prices, and a two-week adoption checklist.
Claude Opus 4.8: the effort dial, fast mode and token math
Claude Opus 4.8 arrives at unchanged pricing with an effort control on all plans, a fast mode at a third of the previous fast-inference cost, and a Messages API change that lets system entries sit inside the messages array so mid-task instruction updates no longer invalidate the prompt cache. Worked token math shows cache hit rate remains the biggest cost lever, and a four-question framework matches effort, speed and fan-out to each workload.