Posts tagged with "Microsoft Foundry"

Found 26 posts

AI & Cloud Infrastructure
September 6, 2026

ChatGPT, Claude and Grok down at once: your failover plan

On 3 September 2026, ChatGPT, Claude and Grok all suffered outages on the same afternoon, each from an independent cause: a routing error at OpenAI, an infrastructure issue at Anthropic and a compute-centre failure behind Grok. A runtime failover plan for Azure teams: Foundry's model router with automatic fallback, APIM's AI gateway and circuit breaker, what redundancy really costs, and the Data Zone and DORA constraints that shape it.

AI Failover
Resilience
Microsoft Foundry
Model Router
Azure API Management
DORA
EU Data Zone
Multi-Model
Outage
By Falak Mahmood
Azure & Cloud
September 4, 2026

GPT-6 Astra in Foundry: the price, the gate and the EU gap

GPT-6 Astra arrived in Microsoft Foundry on 3 September 2026 at $10/$50 per million tokens, the same list price as Claude Fable 5.1 and 2.5x GPT-5.6 Sol on promo, behind a Limited Access gate and with no EU Data Zone. The cost math on document jobs and 60-turn agent loops, the 272K long-context cliff, what the Critical cyber rating means for refusals, and how a Swedish team should handle residency until the EU zone lands at its new 20% premium.

GPT-6 Astra
OpenAI
Microsoft Foundry
Azure OpenAI
AI Agents
LLM Cost
Data Residency
EU AI Act
Limited Access
GPT-5.6
By Falak Mahmood
Azure & Cloud
September 3, 2026

Claude Fable 5.1 in Foundry: cache math and the EU caveats

Claude Fable 5.1 keeps Fable 5's $10/$50 pricing but cuts cache reads to $0.25 per million tokens, which turns a 60-turn agent session from $17.25 into $10.50 and shrinks the premium over Opus 5 from 2x to about 22%. On Microsoft Foundry it ships Anthropic-hosted only, with no EU data zone, no Batches API, a zero default quota on pay-as-you-go, mandatory 30-day retention until Enterprise Frontier Safeguards arrive, and three breaking changes for teams migrating from Fable 5.

Claude Fable 5.1
Anthropic
Microsoft Foundry
Azure
Prompt Caching
AI Agents
Data Residency
GDPR
EU AI Act
LLM Cost
By Falak Mahmood
Business & Strategy
September 2, 2026

OpenAI cuts off Cursor: your model exit plan on Azure

OpenAI will stop supplying models to Cursor on 12 November 2026, invoking a change-of-control clause after SpaceX's $60 billion acquisition, and Cursor absorbed it because OpenAI carried only 5 percent of its traffic. For Azure buyers the lesson is a model exit plan: what Foundry's lifecycle policy and Microsoft's OpenAI licence actually guarantee, where Claude's EU Data Zone gap bites, the cost of a warm second source, and what DORA already requires.

OpenAI
Cursor
SpaceX
Microsoft Foundry
Vendor Risk
Model Portability
DORA
Claude
Azure OpenAI
Exit Strategy
By Falak Mahmood
Azure & Cloud
September 1, 2026

Foundry EU Data Zone premium doubles: the Swedish cost math

From 1 September 2026 Microsoft Foundry charges 20 percent over Global for EU Data Zone deployments, up from 10 percent, and 30 percent for Sweden Central regional, while Global pricing stays flat and West Europe regional hits 50 percent. Pay-as-you-go customers only pay the new rate on models launched from today, PTU customers pay immediately, and GPT-5.6 is not yet on regional Standard in Sweden, so here is the cost math and a six-step checklist.

Microsoft Foundry
Azure OpenAI
EU Data Zone
Data Residency
Sweden Central
GPT-5.6
Provisioned Throughput
Cost Analysis
Azure Policy
GDPR
By Falak Mahmood
AI & Machine Learning
August 31, 2026

Anthropic Claude prompt caching pricing: write, read, TTL math

Anthropic prices prompt caching with three numbers: a 1.25x or 2x premium on cache writes depending on TTL, a 0.1x rate on cache reads, and the base input rate for everything after the last breakpoint. This deep-dive verifies every figure against the current official docs and covers per-model minimums, break-even math, batch stacking and how Claude on Azure converts it all into CCUs.

Prompt Caching
Anthropic
Claude
Claude API
LLM Cost Optimization
Microsoft Foundry
Azure
Batch API
By Falak Mahmood
Azure & Cloud
August 31, 2026

Cohere Parse v5 in Foundry: document parsing cost math

Cohere Parse v5 landed in Microsoft Foundry on 27 August 2026 at $1.50 per 1,000 pages, the same price as Document Intelligence Read and under a third of Content Understanding's layout meter. We price all four Azure document parsers in Sweden Central at 100k, 1M and 10M pages a month, then weigh the saving against Parse's missing confidence scores, absent Swedish language support and preview-only status.

Cohere Parse
Microsoft Foundry
Azure Document Intelligence
Content Understanding
Mistral OCR
Document Parsing
RAG
Cost Analysis
Data Residency
By Falak Mahmood
Azure & Cloud
August 28, 2026

Azure Assistants API retired: migrating to Foundry Agents

The Azure OpenAI Assistants API reached its retirement date on 26 August 2026, and the classic Foundry Agent Service it underpins retires 31 March 2027. A step-by-step migration guide to the new Foundry Agent Service on the Responses API: threads become conversations, runs become responses, assistants become versioned agents, and Microsoft's migration tool rewrites code but not stored state.

Azure OpenAI
Assistants API
Foundry Agent Service
Responses API
Migration
AI Agents
Microsoft Foundry
Azure
GDPR
By Falak Mahmood
AI & Machine Learning
August 14, 2026

Workhorse shootout: Gemini 3.7 Flash, GPT-5.6 Luna, Sonnet 5

Google shipped gemini-3.7-flash as generally available on 13 August 2026 at an introductory $0.75/$3.75 per million tokens, two weeks after OpenAI cut GPT-5.6 Luna by 80 percent and days after Anthropic locked Claude Sonnet 5 at $2/$10 permanently. We compare the three workhorse models on list price, context, Azure availability and EU residency, and show why cost per completed task beats cost per token.

Gemini 3.7 Flash
GPT-5.6 Luna
Claude Sonnet 5
Model Comparison
LLM Pricing
Azure OpenAI
Microsoft Foundry
Enterprise AI
By Falak Mahmood
Security & Compliance
August 13, 2026

Claude's text watermark: what it means for Article 50

Anthropic will weave an invisible watermark into Claude's text output to meet the EU AI Act's Article 50 marking obligation, applying it globally across the API, apps and cloud platforms including Microsoft Foundry. What the mark can and cannot prove, which deployer duties remain yours, and when a DIY provenance layer still earns its keep on Azure.

EU AI Act
Article 50
Claude
Anthropic
Watermarking
C2PA
Microsoft Foundry
Compliance
AI Governance
Content Provenance
By Falak Mahmood
Business & Strategy
August 12, 2026

LLM cost planning autumn 2026: Sonnet 5 stays at $2/$10

Anthropic has cancelled the Claude Sonnet 5 price increase scheduled for 1 September 2026, making the introductory $2 input / $10 output per million tokens the permanent standard price. For teams running Claude on Azure through Microsoft Foundry, that removes a planned 50% jump from autumn budgets and reshapes the mid-tier price comparison against GPT-5.6 Terra and Gemini 3.1 Pro.

LLM Pricing
Claude Sonnet 5
Azure AI Foundry
Microsoft Foundry
GPT-5.6
Gemini
Cost Optimization
AI Budgeting
Anthropic
By Falak Mahmood
Azure & Cloud
August 3, 2026

Azure OpenAI cost check: GPT-5.6 price cuts and PTU math

OpenAI cut GPT-5.6 Luna prices by 80 percent and Terra by 20 percent on 30 July 2026, and Microsoft confirmed the same decreases reach Azure OpenAI Global Standard deployments from 1 August. For Swedish and EU teams running these models on Azure, the cuts move the break-even point for PTU reservations, model routing and residency premiums, so the autumn budget math deserves a fresh pass before any new one-year commitments.

Azure OpenAI
GPT-5.6
Microsoft Foundry
PTU
Provisioned Throughput
Cost Optimization
Pricing
Data Zones
EU Data Residency
By Falak Mahmood
AI & Machine Learning
July 27, 2026

Claude Opus 5 for long-running agents: the cost math

Claude Opus 5 launched on 24 July 2026 at $5/$25 per million tokens with a 1M context window and day-one availability in Microsoft Foundry. For long-running agents the per-token price is the wrong unit: we work through cost per completed task against Sonnet 5 and GPT-5.6 Sol, and flag the EU data-residency caveat Swedish Azure teams need to check first.

Claude Opus 5
Anthropic
AI Agents
Microsoft Foundry
Azure
LLM Pricing
Claude Sonnet 5
GPT-5.6
Cost Optimization
Data Residency
By Falak Mahmood
Azure & Cloud
July 24, 2026

Sovereign AI on Azure: what the Microsoft-Mistral deal means

Microsoft and Mistral announced an expanded partnership on 21 July 2026: Mistral Medium 3.5 and OCR 4 arrive in Microsoft Foundry and Copilot Studio, deployable from Azure cloud to customer-controlled and fully air-gapped Azure Local environments, backed by a multibillion-dollar European GPU buildout. We compare the three deployment modes and what each one solves for Swedish public sector and regulated industries.

Sovereign AI
Microsoft Foundry
Mistral
Azure Local
Copilot Studio
Data Residency
EU AI Act
Public Sector
Regulated Industries
By Falak Mahmood
AI & Cloud Infrastructure
July 15, 2026

Running GPT-5.6 the enterprise way on Microsoft Foundry

GPT-5.6 (Sol, Terra, Luna) went GA in Microsoft Foundry on 9 July 2026, day-and-date with OpenAI, alongside a new Asia-Pacific Data Zone and a hosted agents runtime with VNet integration. A practical guide for Swedish and EU Azure teams: choosing between the three models, picking Global Standard versus EU Data Zone versus PTUs, worked cost math on the launch prices, and a two-week adoption checklist.

GPT-5.6
Microsoft Foundry
Azure OpenAI
Data Zones
EU Data Residency
Hosted Agents
VNet
AI Agents
LLM Pricing
By Falak Mahmood
AI & Machine Learning
July 13, 2026

GPT-5.6 Sol vs Terra vs Luna: an Azure routing playbook

OpenAI released the GPT-5.6 series on 9 July 2026 in three tiers: Sol for hard reasoning and long autonomous runs, Terra for everyday work and Luna for speed and cost, with same-day availability in Microsoft Foundry and a new preferred-model role in Microsoft 365 Copilot. This playbook maps Azure workloads to the right tier, works the token math in SEK and flags the Copilot subprocessor setting Swedish admins must review before 24 July.

GPT-5.6
OpenAI
Azure OpenAI
Microsoft Foundry
Model Routing
LLM Cost Optimization
Microsoft 365 Copilot
ChatGPT Work
EU Data Residency
By Falak Mahmood
AI & Machine Learning
July 1, 2026

Claude Sonnet 5 vs Opus 4.8 vs GPT-5.5: agent cost math

Anthropic launched Claude Sonnet 5 on 30 June 2026 at an introductory 2/10 dollars per million tokens, posting 63.2% on SWE-bench Pro and near-Opus agentic performance at 40-60% of the cost per task. We run the cost math against Opus 4.8 and GPT-5.5, set out a routing framework for when the cheap model wins, and draw the continuity lesson from the eighteen-day Fable and Mythos export-control pause.

Claude Sonnet 5
Claude Opus 4.8
GPT-5.5
Anthropic
AI Agents
Model Selection
Cost Optimization
Microsoft Foundry
Azure
By Falak Mahmood
AI & Cloud Infrastructure
June 30, 2026

Claude on Azure is GA: Foundry deployment and CCU costs

Claude Opus 4.8 and Claude Haiku 4.5 are now generally available in Microsoft Foundry, hosted on Azure with Entra ID authentication, prompt caching, extended thinking and billing through Claude Consumption Units on your existing Azure invoice. Deployment steps, the CCU cost model compared with Azure OpenAI, and the data-residency caveats Swedish and EU teams should assess before production use.

Claude
Microsoft Foundry
Azure
Anthropic
Claude Consumption Units
Azure OpenAI
Data Residency
GDPR
Entra ID
By Falak Mahmood
AI & Cloud Infrastructure
June 3, 2026

Build 2026 Foundry agents: what Azure teams can ship now

Microsoft Build 2026 turned Foundry into a full production-agent stack: Foundry IQ for unified retrieval, Toolboxes for managed tool access, agent memory, Voice Live and the experimental Scout Autopilot. Foundry IQ knowledge bases and Voice Live are generally available now, Toolboxes and memory sit in public preview, and Scout remains experimental, which sets the build, pilot and watch lanes for an Azure-first EU team.

Microsoft Build 2026
Microsoft Foundry
AI Agents
Foundry IQ
Azure
RAG
Agent Memory
EU Compliance
By Falak Mahmood
Security & Compliance
April 14, 2026

Azure Entra Agent ID: Identity and Permissions for Agentic AI

A deep dive into Microsoft Entra Agent ID, the control plane for AI agent identity in 2026. Covers identity blueprints, attended and unattended authentication, tool-level RBAC, conditional access, OBO flows across multi-agent systems, and the audit logging that satisfies DORA, NIS2, and AI Act obligations.

Entra
Azure AD
Identity
AI Agents
Agentic AI
RBAC
Zero Trust
Security
Microsoft Foundry
Compliance
By Falak Mahmood
Microsoft Ignite 2025
November 28, 2025

Training and Deploying Custom Reasoning Models with Azure ML and Foundry - Microsoft Ignite 2025

See the magic happen in real time. Learn how to train and deploy custom reasoning models with Azure ML and Microsoft Foundry—from fine-tuning to reinforcement learning, performance optimization with speculative decoding, distillation, and production deployment delivering measurable ROI.

Microsoft Ignite
Azure Machine Learning
Microsoft Foundry
Custom Models
Fine-Tuning
Reinforcement Learning
Model Optimization
Speculative Decoding
Model Distillation
Quantization
Kubernetes
AI Training
Model Deployment
Performance Optimization
Enterprise AI
By Falak Mahmood
AI & Cloud Infrastructure
November 28, 2025

Autonomous Agents Powered by Reasoning Models: Building Intelligent AI with Microsoft Foundry - Microsoft Ignite 2025

Microsoft Ignite BRK203: Reasoning models as the brains behind autonomous agents. Multi-step problem solving, explainable decisions, agentic workflows (lead scoring, content generation, support). Foundry 11,000+ model catalog, customer stories from healthcare and legal sectors.

Microsoft Ignite 2025
Reasoning Models
Autonomous Agents
Microsoft Foundry
OpenAI o1
OpenAI o3
Claude 3.5 Sonnet
Agentic AI
Chain-of-Thought
Lead Scoring
Content Generation
Customer Support AI
Healthcare AI
Legal Tech AI
Multi-Agent Systems
Model Routing
Parallel Function Calling
Explainable AI
Azure OpenAI
By Falak Mahmood
Microsoft Ignite 2025
November 28, 2025

Microsoft Foundry: The AI Platform for the Agentic Era - Ignite 2025

From scientific research to enterprise AI transformation, discover how Microsoft Foundry unifies models from OpenAI, Anthropic, Cohere, Meta, and more into one secure platform. Learn intelligent model routing, cost optimization, and the game-changing Claude integration.

Microsoft Ignite
Microsoft Foundry
Azure AI
Anthropic Claude
Multi-Model AI
AI Agents
OpenAI
Cohere
Meta Llama
Enterprise AI
AI Platform
Model Orchestration
Intelligent Routing
Cost Optimization
AI Security
Responsible AI
Agentic AI
By Falak Mahmood
AI & Cloud Infrastructure
November 28, 2025

Model Context Protocol: The Future of Agent-Tool Interactions - Microsoft Ignite 2025

Microsoft Ignite BRK194: Model Context Protocol (MCP) standardizes agent-tool communication across platforms. Azure API Center integration, federated registries, cross-cloud orchestration, and enterprise governance for scalable agentic ecosystems.

Microsoft Ignite 2025
Model Context Protocol
MCP
Microsoft Foundry
Azure API Center
AI Agents
Agent Orchestration
Enterprise AI
Tool Integration
Federated Registries
Microsoft Entra
GitHub Copilot
VS Code
Logic Apps
Databricks
Cross-Platform AI
By Falak Mahmood
Microsoft Ignite 2025
November 28, 2025

Microsoft Foundry: The Enterprise Agent Factory - Microsoft Ignite 2025

Ride the agent revolution with Microsoft Foundry, the enterprise-ready Agent Factory. Build, test, and launch intelligent agents with 1,400+ tools, 11,000+ models, multi-agent orchestration, and seamless Microsoft 365 integration—all with bulletproof security and governance.

Microsoft Ignite
Microsoft Foundry
Azure AI Foundry
AI Agents
Agent Factory
Foundry IQ
Multi-Agent Orchestration
Microsoft Agent Framework
Copilot Studio
Agent Memory
Synthetic Data
Model Fine-Tuning
Enterprise AI
Microsoft 365 Integration
Agent Governance
By Falak Mahmood
AI & Cloud Infrastructure
November 28, 2025

Fine-Tuning in Microsoft Foundry: Building Production-Ready AI Agents - Microsoft Ignite 2025

Microsoft Ignite BRK188: Fine-tuning in Microsoft Foundry transforms generic models into production-ready agents. Synthetic data generation, supervised + reinforcement fine-tuning, 40-90% cost reduction, 95%+ accuracy. Real-world results: 2M docs/day, $27M savings.

Microsoft Ignite 2025
Microsoft Foundry
Fine-Tuning
Supervised Fine-Tuning
Reinforcement Fine-Tuning
Agentic RFT
Synthetic Data Generation
Azure OpenAI
Tool Calling
Data Extraction
Workflow Execution
Model Optimization
Production AI
Agent Accuracy
Cost Optimization
GPT-4o
By Falak Mahmood