Posts tagged with "Azure OpenAI"

Found 21 posts

Azure & Cloud
September 4, 2026

GPT-6 Astra in Foundry: the price, the gate and the EU gap

GPT-6 Astra arrived in Microsoft Foundry on 3 September 2026 at $10/$50 per million tokens, the same list price as Claude Fable 5.1 and 2.5x GPT-5.6 Sol on promo, behind a Limited Access gate and with no EU Data Zone. The cost math on document jobs and 60-turn agent loops, the 272K long-context cliff, what the Critical cyber rating means for refusals, and how a Swedish team should handle residency until the EU zone lands at its new 20% premium.

GPT-6 Astra
OpenAI
Microsoft Foundry
Azure OpenAI
AI Agents
LLM Cost
Data Residency
EU AI Act
Limited Access
GPT-5.6
By Falak Mahmood
Business & Strategy
September 2, 2026

OpenAI cuts off Cursor: your model exit plan on Azure

OpenAI will stop supplying models to Cursor on 12 November 2026, invoking a change-of-control clause after SpaceX's $60 billion acquisition, and Cursor absorbed it because OpenAI carried only 5 percent of its traffic. For Azure buyers the lesson is a model exit plan: what Foundry's lifecycle policy and Microsoft's OpenAI licence actually guarantee, where Claude's EU Data Zone gap bites, the cost of a warm second source, and what DORA already requires.

OpenAI
Cursor
SpaceX
Microsoft Foundry
Vendor Risk
Model Portability
DORA
Claude
Azure OpenAI
Exit Strategy
By Falak Mahmood
Azure & Cloud
September 1, 2026

Foundry EU Data Zone premium doubles: the Swedish cost math

From 1 September 2026 Microsoft Foundry charges 20 percent over Global for EU Data Zone deployments, up from 10 percent, and 30 percent for Sweden Central regional, while Global pricing stays flat and West Europe regional hits 50 percent. Pay-as-you-go customers only pay the new rate on models launched from today, PTU customers pay immediately, and GPT-5.6 is not yet on regional Standard in Sweden, so here is the cost math and a six-step checklist.

Microsoft Foundry
Azure OpenAI
EU Data Zone
Data Residency
Sweden Central
GPT-5.6
Provisioned Throughput
Cost Analysis
Azure Policy
GDPR
By Falak Mahmood
Azure & Cloud
August 28, 2026

Azure Assistants API retired: migrating to Foundry Agents

The Azure OpenAI Assistants API reached its retirement date on 26 August 2026, and the classic Foundry Agent Service it underpins retires 31 March 2027. A step-by-step migration guide to the new Foundry Agent Service on the Responses API: threads become conversations, runs become responses, assistants become versioned agents, and Microsoft's migration tool rewrites code but not stored state.

Azure OpenAI
Assistants API
Foundry Agent Service
Responses API
Migration
AI Agents
Microsoft Foundry
Azure
GDPR
By Falak Mahmood
AI & Machine Learning
August 14, 2026

Workhorse shootout: Gemini 3.7 Flash, GPT-5.6 Luna, Sonnet 5

Google shipped gemini-3.7-flash as generally available on 13 August 2026 at an introductory $0.75/$3.75 per million tokens, two weeks after OpenAI cut GPT-5.6 Luna by 80 percent and days after Anthropic locked Claude Sonnet 5 at $2/$10 permanently. We compare the three workhorse models on list price, context, Azure availability and EU residency, and show why cost per completed task beats cost per token.

Gemini 3.7 Flash
GPT-5.6 Luna
Claude Sonnet 5
Model Comparison
LLM Pricing
Azure OpenAI
Microsoft Foundry
Enterprise AI
By Falak Mahmood
Business & Strategy
August 10, 2026

Unlimited free ChatGPT vs governed enterprise AI on Azure

OpenAI removed limits on text chats for free ChatGPT users on 6 August 2026 and made GPT-5.6 Luna the default, cutting factual errors by roughly 62 percent versus the prior model. For Swedish and EU enterprises on Azure, the free consumer tool employees already use just became unlimited and much stronger, so shadow AI pressure rises and the case for a governed answer built on Copilot Chat, paid Copilot seats and Azure OpenAI becomes urgent.

Shadow AI
ChatGPT
GPT-5.6
Microsoft 365 Copilot
Azure OpenAI
AI Governance
EU AI Act
GDPR
Enterprise AI Strategy
By Falak Mahmood
Security & Compliance
August 4, 2026

AI Act enforcement is now real: an Azure deployer checklist

On 2 August 2026 the European Commission's enforcement powers over general-purpose AI providers activated: the AI Office can now demand documentation, run model evaluations, restrict models from the EU market and fine up to 3% of global turnover or EUR 15 million. The same date brought Article 50 transparency into application, and this guide maps what Azure OpenAI and Foundry teams must demand from vendors versus handle themselves as deployers.

AI Act
EU AI Act
GPAI
Article 50
Azure OpenAI
Azure AI Foundry
AI Office
Compliance
Digital Omnibus
Transparency
By Falak Mahmood
Azure & Cloud
August 3, 2026

Azure OpenAI cost check: GPT-5.6 price cuts and PTU math

OpenAI cut GPT-5.6 Luna prices by 80 percent and Terra by 20 percent on 30 July 2026, and Microsoft confirmed the same decreases reach Azure OpenAI Global Standard deployments from 1 August. For Swedish and EU teams running these models on Azure, the cuts move the break-even point for PTU reservations, model routing and residency premiums, so the autumn budget math deserves a fresh pass before any new one-year commitments.

Azure OpenAI
GPT-5.6
Microsoft Foundry
PTU
Provisioned Throughput
Cost Optimization
Pricing
Data Zones
EU Data Residency
By Falak Mahmood
AI & Machine Learning
July 23, 2026

Tokens are the new pricing lever: Gemini 3.6 Flash math

Google's 21 July release of Gemini 3.6 Flash pairs an output-price cut from $9.00 to $7.50 per million tokens with a claim of roughly 17% fewer output tokens per task, compounding to about 31% lower output cost for unchanged work. That combination makes per-million-token price sheets unreliable for model comparison, and Azure teams should measure cost per completed task instead.

Gemini
Google
LLM Pricing
Token Efficiency
Azure OpenAI
Azure AI Foundry
FinOps
Cost Optimization
AI Procurement
By Falak Mahmood
Security & Compliance
July 22, 2026

Article 50 compliance for Azure OpenAI apps: a guide

The European Commission adopted its final Article 50 transparency guidelines on 20 July 2026 and confirmed the Code of Practice on marking AI-generated content as adequate, less than two weeks before the obligations start to apply. Here is what Swedish and EU teams running chatbots, copilots and content generators on Azure OpenAI must implement: chatbot disclosure, machine-readable marking and deepfake labels, with concrete code patterns for each.

EU AI Act
Article 50
Azure OpenAI
Transparency
C2PA
Code of Practice
Compliance
GenAI
Chatbots
By Falak Mahmood
AI & Cloud Infrastructure
July 15, 2026

Running GPT-5.6 the enterprise way on Microsoft Foundry

GPT-5.6 (Sol, Terra, Luna) went GA in Microsoft Foundry on 9 July 2026, day-and-date with OpenAI, alongside a new Asia-Pacific Data Zone and a hosted agents runtime with VNet integration. A practical guide for Swedish and EU Azure teams: choosing between the three models, picking Global Standard versus EU Data Zone versus PTUs, worked cost math on the launch prices, and a two-week adoption checklist.

GPT-5.6
Microsoft Foundry
Azure OpenAI
Data Zones
EU Data Residency
Hosted Agents
VNet
AI Agents
LLM Pricing
By Falak Mahmood
AI & Machine Learning
July 13, 2026

GPT-5.6 Sol vs Terra vs Luna: an Azure routing playbook

OpenAI released the GPT-5.6 series on 9 July 2026 in three tiers: Sol for hard reasoning and long autonomous runs, Terra for everyday work and Luna for speed and cost, with same-day availability in Microsoft Foundry and a new preferred-model role in Microsoft 365 Copilot. This playbook maps Azure workloads to the right tier, works the token math in SEK and flags the Copilot subprocessor setting Swedish admins must review before 24 July.

GPT-5.6
OpenAI
Azure OpenAI
Microsoft Foundry
Model Routing
LLM Cost Optimization
Microsoft 365 Copilot
ChatGPT Work
EU Data Residency
By Falak Mahmood
AI & Cloud Infrastructure
June 30, 2026

Claude on Azure is GA: Foundry deployment and CCU costs

Claude Opus 4.8 and Claude Haiku 4.5 are now generally available in Microsoft Foundry, hosted on Azure with Entra ID authentication, prompt caching, extended thinking and billing through Claude Consumption Units on your existing Azure invoice. Deployment steps, the CCU cost model compared with Azure OpenAI, and the data-residency caveats Swedish and EU teams should assess before production use.

Claude
Microsoft Foundry
Azure
Anthropic
Claude Consumption Units
Azure OpenAI
Data Residency
GDPR
Entra ID
By Falak Mahmood
Security & Compliance
June 17, 2026

AI Act deadlines moved: what still lands August 2, 2026

On 16 June 2026 the European Parliament approved the Digital Omnibus amendments 423-57, moving Annex III high-risk AI Act obligations to 2 December 2027 and product-embedded obligations to 2 August 2028. Article 50 transparency duties and the Commission's GPAI enforcement powers were not delayed, which leaves Swedish enterprises six weeks to ship chatbot disclosure, content marking and a documented GPAI position before 2 August 2026.

EU AI Act
Digital Omnibus
Article 50
GPAI
Compliance
AI Governance
Azure OpenAI
Sweden
Transparency
By Falak Mahmood
AI & Cloud Infrastructure
May 15, 2026

Prompt Caching in 2026: Cut Azure OpenAI and Claude Costs

Prompt caching is the highest-ROI cost lever on long-context LLM workloads in 2026. Anthropic, OpenAI, and Azure OpenAI all offer it with different pricing and breakpoint semantics. A worked comparison of the three providers, the placement patterns that actually hit cache, where the cache silently goes cold, and a 30-minute audit that pays back.

Prompt Caching
Cost Optimization
Anthropic
OpenAI
Azure OpenAI
By Falak Mahmood
AI & Cloud Infrastructure
April 9, 2026

LLM vs AI Agent vs Agentic AI: Drawing the Lines That Matter

The capability spectrum from stateless LLM to multi-agent orchestration is one of the most conflated concepts in the 2026 AI market. The distinctions matter. They change architecture, they change cost by an order of magnitude, and under the EU AI Act they change compliance posture.

AI Agents
Agentic AI
LLM
AI Architecture
Azure OpenAI
Claude
Multi-Agent Systems
Tool Use
By Falak Mahmood
AI & Cloud Infrastructure
April 2, 2026

Cost-Optimizing Azure OpenAI: PTUs, Batch, Caching in 2026

A concrete playbook for reducing Azure OpenAI bills in 2026. Break-even math for Provisioned Throughput Units, prompt-cache economics, the Batch API 50 percent discount, Foundry IQ for retrieval, tiered model routing, and the telemetry that keeps the wins honest.

Azure OpenAI
Cost Optimization
PTU
Foundry IQ
LLM
By Falak Mahmood
AI & Cloud Infrastructure
March 21, 2026

RAG for Manufacturing: Grounding LLMs in Technical Docs

Generic LLM copilots are a liability in manufacturing. Technicians need answers that cite the exact procedure, not plausible-sounding text. Retrieval-augmented generation grounded in Azure AI Search solves this when architected correctly. This is the pattern that holds up under service-bay pressure.

RAG
Manufacturing
LLM
Azure OpenAI
Azure AI Search
By Falak Mahmood
AI & Cloud Infrastructure
November 28, 2025

Autonomous Agents Powered by Reasoning Models: Building Intelligent AI with Microsoft Foundry - Microsoft Ignite 2025

Microsoft Ignite BRK203: Reasoning models as the brains behind autonomous agents. Multi-step problem solving, explainable decisions, agentic workflows (lead scoring, content generation, support). Foundry 11,000+ model catalog, customer stories from healthcare and legal sectors.

Microsoft Ignite 2025
Reasoning Models
Autonomous Agents
Microsoft Foundry
OpenAI o1
OpenAI o3
Claude 3.5 Sonnet
Agentic AI
Chain-of-Thought
Lead Scoring
Content Generation
Customer Support AI
Healthcare AI
Legal Tech AI
Multi-Agent Systems
Model Routing
Parallel Function Calling
Explainable AI
Azure OpenAI
By Falak Mahmood
AI & Cloud Infrastructure
November 28, 2025

Building Knowledge-Powered Agents with Azure AI Search: RAG, Hybrid Search, and Agentic Retrieval - Microsoft Ignite 2025

Microsoft Ignite BRK193: Build agents with Azure AI Search knowledge features. Connect to SharePoint, web, blob. Hybrid search (keyword+vector+semantic), agentic retrieval with query planning, reasoning effort modes, Foundry IQ with MCP protocol. Code-focused implementation guide.

Microsoft Ignite 2025
Azure AI Search
RAG
Retrieval-Augmented Generation
Agentic Retrieval
Hybrid Search
Vector Search
Semantic Ranking
SharePoint Integration
Knowledge Agents
Query Planning
Reasoning Effort Modes
Foundry IQ
MCP Protocol
Azure OpenAI
Document Indexing
Reciprocal Rank Fusion
Web Crawler
By Falak Mahmood
AI & Cloud Infrastructure
November 28, 2025

Fine-Tuning in Microsoft Foundry: Building Production-Ready AI Agents - Microsoft Ignite 2025

Microsoft Ignite BRK188: Fine-tuning in Microsoft Foundry transforms generic models into production-ready agents. Synthetic data generation, supervised + reinforcement fine-tuning, 40-90% cost reduction, 95%+ accuracy. Real-world results: 2M docs/day, $27M savings.

Microsoft Ignite 2025
Microsoft Foundry
Fine-Tuning
Supervised Fine-Tuning
Reinforcement Fine-Tuning
Agentic RFT
Synthetic Data Generation
Azure OpenAI
Tool Calling
Data Extraction
Workflow Execution
Model Optimization
Production AI
Agent Accuracy
Cost Optimization
GPT-4o
By Falak Mahmood