GPT-6 Astra in Foundry: the price, the gate and the EU gap
GPT-6 Astra arrived in Microsoft Foundry on 3 September 2026 at $10/$50 per million tokens, the same list price as Claude Fable 5.1 and 2.5x GPT-5.6 Sol on promo, behind a Limited Access gate and with no EU Data Zone. The cost math on document jobs and 60-turn agent loops, the 272K long-context cliff, what the Critical cyber rating means for refusals, and how a Swedish team should handle residency until the EU zone lands at its new 20% premium.
OpenAI cuts off Cursor: your model exit plan on Azure
OpenAI will stop supplying models to Cursor on 12 November 2026, invoking a change-of-control clause after SpaceX's $60 billion acquisition, and Cursor absorbed it because OpenAI carried only 5 percent of its traffic. For Azure buyers the lesson is a model exit plan: what Foundry's lifecycle policy and Microsoft's OpenAI licence actually guarantee, where Claude's EU Data Zone gap bites, the cost of a warm second source, and what DORA already requires.
GPT-5.6 Sol vs Terra vs Luna: an Azure routing playbook
OpenAI released the GPT-5.6 series on 9 July 2026 in three tiers: Sol for hard reasoning and long autonomous runs, Terra for everyday work and Luna for speed and cost, with same-day availability in Microsoft Foundry and a new preferred-model role in Microsoft 365 Copilot. This playbook maps Azure workloads to the right tier, works the token math in SEK and flags the Copilot subprocessor setting Swedish admins must review before 24 July.
Prompt Caching in 2026: Cut Azure OpenAI and Claude Costs
Prompt caching is the highest-ROI cost lever on long-context LLM workloads in 2026. Anthropic, OpenAI, and Azure OpenAI all offer it with different pricing and breakpoint semantics. A worked comparison of the three providers, the placement patterns that actually hit cache, where the cache silently goes cold, and a 30-minute audit that pays back.
Building Reliable Agent Tools: Schemas, Idempotency, Recovery
A production-shaped guide to designing AI agent tools that the model can actually use without breaking things. Schema choices, idempotency keys, error responses the model can act on, granularity tradeoffs, versioning, and the patterns that separate demo-quality tools from ones that hold up in real workloads.
Prompt Caching: Cutting LLM Costs Without Quality Loss
A technical guide to prompt caching across Claude, Azure OpenAI, and GPT — what belongs in the cache, how to structure cache breakpoints, TTL realities, hit-rate optimization, and the anti-patterns that erase the savings.
Microsoft Foundry: The AI Platform for the Agentic Era - Ignite 2025
From scientific research to enterprise AI transformation, discover how Microsoft Foundry unifies models from OpenAI, Anthropic, Cohere, Meta, and more into one secure platform. Learn intelligent model routing, cost optimization, and the game-changing Claude integration.