Blog

Technical depth on AI, Azure, Next.js, and the engineering decisions behind them.

244 posts

MCP goes stateless: Azure MCP server migration checklist

MCP's 2026-07-28 release candidate removes the initialize handshake and the Mcp-Session-Id header, so every request is self-contained and Azure-hosted MCP servers can run behind plain load balancers without sticky sessions or Redis session stores. Tasks and MCP Apps land as formal extensions, six SEPs harden authorization around OAuth 2.0 and OpenID Connect, and Roots, Sampling and Logging enter a 12-month deprecation window.

May 25, 2026149 viewsAI & Cloud Infrastructure

Gemini Enterprise Agent Platform vs Azure AI Foundry

Google's I/O 2026 enterprise announcements, led by the Managed Agents API, Gemini Spark and Gemini 3.5 Flash, take direct aim at Azure AI Foundry Agent Service and the Copilot ecosystem, down to launch connectors for SharePoint and OneDrive. For Azure-first Swedish and EU teams the decision rests on four questions: where the data lives, whether the cost claim survives real traces, whether audit obligations can be met, and what a second platform costs.

May 22, 2026144 viewsAI & Cloud Infrastructure

Gemini 3.5 Flash vs GPT-5.5 vs Claude Opus 4.7 at I/O 2026

Google shipped Gemini 3.5 Flash on day one of I/O 2026, the third frontier agentic coding model in five weeks after Claude Opus 4.7 and GPT-5.5, alongside a $100/month developer tier and the Antigravity 2.0 agent platform. For Azure-first teams the practical response is a reusable evaluation harness and a documented cost comparison for renewal leverage, not a migration.

May 20, 2026127 viewsAI & Machine Learning

Small Models in Production: When Phi-4 and 8B Llama Win

Frontier models are the default. Defaults are how teams overpay on LLM bills. Three workloads where small models (Phi-4, Llama 3.x 8B, Mistral Small) outperform on cost-per-decision without losing meaningfully on quality, three workloads where they do not, and the two-tier production pattern that cost-conscious teams converge on after a quarter of evaluation work.

May 18, 2026887 viewsAI & Cloud Infrastructure

Prompt Caching in 2026: Cut Azure OpenAI and Claude Costs

Prompt caching is the highest-ROI cost lever on long-context LLM workloads in 2026. Anthropic, OpenAI, and Azure OpenAI all offer it with different pricing and breakpoint semantics. A worked comparison of the three providers, the placement patterns that actually hit cache, where the cache silently goes cold, and a 30-minute audit that pays back.

May 15, 20269,975 viewsAI & Cloud Infrastructure

EU AI Act High-Risk Readiness: 11 Weeks to August 2026

High-risk obligations under the EU AI Act apply August 2, 2026. For teams shipping AI inside the Annex III categories, that is 11 weeks of runway. This is the readiness state most teams reach by skipping the strategy decks: what actually changes, what the four obligations that require code look like, and where compliance spending misfires before the deadline.

May 10, 2026618 viewsSecurity & Compliance