GPT-6 Astra in Foundry: the price, the gate and the EU gap
GPT-6 Astra arrived in Microsoft Foundry on 3 September 2026 at $10/$50 per million tokens, the same list price as Claude Fable 5.1 and 2.5x GPT-5.6 Sol on promo, behind a Limited Access gate and with no EU Data Zone. The cost math on document jobs and 60-turn agent loops, the 272K long-context cliff, what the Critical cyber rating means for refusals, and how a Swedish team should handle residency until the EU zone lands at its new 20% premium.
OpenAI cuts off Cursor: your model exit plan on Azure
OpenAI will stop supplying models to Cursor on 12 November 2026, invoking a change-of-control clause after SpaceX's $60 billion acquisition, and Cursor absorbed it because OpenAI carried only 5 percent of its traffic. For Azure buyers the lesson is a model exit plan: what Foundry's lifecycle policy and Microsoft's OpenAI licence actually guarantee, where Claude's EU Data Zone gap bites, the cost of a warm second source, and what DORA already requires.
Foundry EU Data Zone premium doubles: the Swedish cost math
From 1 September 2026 Microsoft Foundry charges 20 percent over Global for EU Data Zone deployments, up from 10 percent, and 30 percent for Sweden Central regional, while Global pricing stays flat and West Europe regional hits 50 percent. Pay-as-you-go customers only pay the new rate on models launched from today, PTU customers pay immediately, and GPT-5.6 is not yet on regional Standard in Sweden, so here is the cost math and a six-step checklist.
Azure Assistants API retired: migrating to Foundry Agents
The Azure OpenAI Assistants API reached its retirement date on 26 August 2026, and the classic Foundry Agent Service it underpins retires 31 March 2027. A step-by-step migration guide to the new Foundry Agent Service on the Responses API: threads become conversations, runs become responses, assistants become versioned agents, and Microsoft's migration tool rewrites code but not stored state.
Workhorse shootout: Gemini 3.7 Flash, GPT-5.6 Luna, Sonnet 5
Google shipped gemini-3.7-flash as generally available on 13 August 2026 at an introductory $0.75/$3.75 per million tokens, two weeks after OpenAI cut GPT-5.6 Luna by 80 percent and days after Anthropic locked Claude Sonnet 5 at $2/$10 permanently. We compare the three workhorse models on list price, context, Azure availability and EU residency, and show why cost per completed task beats cost per token.
Unlimited free ChatGPT vs governed enterprise AI on Azure
OpenAI removed limits on text chats for free ChatGPT users on 6 August 2026 and made GPT-5.6 Luna the default, cutting factual errors by roughly 62 percent versus the prior model. For Swedish and EU enterprises on Azure, the free consumer tool employees already use just became unlimited and much stronger, so shadow AI pressure rises and the case for a governed answer built on Copilot Chat, paid Copilot seats and Azure OpenAI becomes urgent.
AI Act enforcement is now real: an Azure deployer checklist
On 2 August 2026 the European Commission's enforcement powers over general-purpose AI providers activated: the AI Office can now demand documentation, run model evaluations, restrict models from the EU market and fine up to 3% of global turnover or EUR 15 million. The same date brought Article 50 transparency into application, and this guide maps what Azure OpenAI and Foundry teams must demand from vendors versus handle themselves as deployers.
Azure OpenAI cost check: GPT-5.6 price cuts and PTU math
OpenAI cut GPT-5.6 Luna prices by 80 percent and Terra by 20 percent on 30 July 2026, and Microsoft confirmed the same decreases reach Azure OpenAI Global Standard deployments from 1 August. For Swedish and EU teams running these models on Azure, the cuts move the break-even point for PTU reservations, model routing and residency premiums, so the autumn budget math deserves a fresh pass before any new one-year commitments.
Tokens are the new pricing lever: Gemini 3.6 Flash math
Google's 21 July release of Gemini 3.6 Flash pairs an output-price cut from $9.00 to $7.50 per million tokens with a claim of roughly 17% fewer output tokens per task, compounding to about 31% lower output cost for unchanged work. That combination makes per-million-token price sheets unreliable for model comparison, and Azure teams should measure cost per completed task instead.
Article 50 compliance for Azure OpenAI apps: a guide
The European Commission adopted its final Article 50 transparency guidelines on 20 July 2026 and confirmed the Code of Practice on marking AI-generated content as adequate, less than two weeks before the obligations start to apply. Here is what Swedish and EU teams running chatbots, copilots and content generators on Azure OpenAI must implement: chatbot disclosure, machine-readable marking and deepfake labels, with concrete code patterns for each.
Running GPT-5.6 the enterprise way on Microsoft Foundry
GPT-5.6 (Sol, Terra, Luna) went GA in Microsoft Foundry on 9 July 2026, day-and-date with OpenAI, alongside a new Asia-Pacific Data Zone and a hosted agents runtime with VNet integration. A practical guide for Swedish and EU Azure teams: choosing between the three models, picking Global Standard versus EU Data Zone versus PTUs, worked cost math on the launch prices, and a two-week adoption checklist.
GPT-5.6 Sol vs Terra vs Luna: an Azure routing playbook
OpenAI released the GPT-5.6 series on 9 July 2026 in three tiers: Sol for hard reasoning and long autonomous runs, Terra for everyday work and Luna for speed and cost, with same-day availability in Microsoft Foundry and a new preferred-model role in Microsoft 365 Copilot. This playbook maps Azure workloads to the right tier, works the token math in SEK and flags the Copilot subprocessor setting Swedish admins must review before 24 July.
Claude on Azure is GA: Foundry deployment and CCU costs
Claude Opus 4.8 and Claude Haiku 4.5 are now generally available in Microsoft Foundry, hosted on Azure with Entra ID authentication, prompt caching, extended thinking and billing through Claude Consumption Units on your existing Azure invoice. Deployment steps, the CCU cost model compared with Azure OpenAI, and the data-residency caveats Swedish and EU teams should assess before production use.
AI Act deadlines moved: what still lands August 2, 2026
On 16 June 2026 the European Parliament approved the Digital Omnibus amendments 423-57, moving Annex III high-risk AI Act obligations to 2 December 2027 and product-embedded obligations to 2 August 2028. Article 50 transparency duties and the Commission's GPAI enforcement powers were not delayed, which leaves Swedish enterprises six weeks to ship chatbot disclosure, content marking and a documented GPAI position before 2 August 2026.
Prompt Caching in 2026: Cut Azure OpenAI and Claude Costs
Prompt caching is the highest-ROI cost lever on long-context LLM workloads in 2026. Anthropic, OpenAI, and Azure OpenAI all offer it with different pricing and breakpoint semantics. A worked comparison of the three providers, the placement patterns that actually hit cache, where the cache silently goes cold, and a 30-minute audit that pays back.
LLM vs AI Agent vs Agentic AI: Drawing the Lines That Matter
The capability spectrum from stateless LLM to multi-agent orchestration is one of the most conflated concepts in the 2026 AI market. The distinctions matter. They change architecture, they change cost by an order of magnitude, and under the EU AI Act they change compliance posture.
Cost-Optimizing Azure OpenAI: PTUs, Batch, Caching in 2026
A concrete playbook for reducing Azure OpenAI bills in 2026. Break-even math for Provisioned Throughput Units, prompt-cache economics, the Batch API 50 percent discount, Foundry IQ for retrieval, tiered model routing, and the telemetry that keeps the wins honest.
RAG for Manufacturing: Grounding LLMs in Technical Docs
Generic LLM copilots are a liability in manufacturing. Technicians need answers that cite the exact procedure, not plausible-sounding text. Retrieval-augmented generation grounded in Azure AI Search solves this when architected correctly. This is the pattern that holds up under service-bay pressure.
Autonomous Agents Powered by Reasoning Models: Building Intelligent AI with Microsoft Foundry - Microsoft Ignite 2025
Microsoft Ignite BRK203: Reasoning models as the brains behind autonomous agents. Multi-step problem solving, explainable decisions, agentic workflows (lead scoring, content generation, support). Foundry 11,000+ model catalog, customer stories from healthcare and legal sectors.
Building Knowledge-Powered Agents with Azure AI Search: RAG, Hybrid Search, and Agentic Retrieval - Microsoft Ignite 2025
Microsoft Ignite BRK193: Build agents with Azure AI Search knowledge features. Connect to SharePoint, web, blob. Hybrid search (keyword+vector+semantic), agentic retrieval with query planning, reasoning effort modes, Foundry IQ with MCP protocol. Code-focused implementation guide.
Fine-Tuning in Microsoft Foundry: Building Production-Ready AI Agents - Microsoft Ignite 2025
Microsoft Ignite BRK188: Fine-tuning in Microsoft Foundry transforms generic models into production-ready agents. Synthetic data generation, supervised + reinforcement fine-tuning, 40-90% cost reduction, 95%+ accuracy. Real-world results: 2M docs/day, $27M savings.