Posts tagged with "LLM Cost"

Found 4 posts

Azure & Cloud
September 4, 2026

GPT-6 Astra in Foundry: the price, the gate and the EU gap

GPT-6 Astra arrived in Microsoft Foundry on 3 September 2026 at $10/$50 per million tokens, the same list price as Claude Fable 5.1 and 2.5x GPT-5.6 Sol on promo, behind a Limited Access gate and with no EU Data Zone. The cost math on document jobs and 60-turn agent loops, the 272K long-context cliff, what the Critical cyber rating means for refusals, and how a Swedish team should handle residency until the EU zone lands at its new 20% premium.

GPT-6 Astra
OpenAI
Microsoft Foundry
Azure OpenAI
AI Agents
LLM Cost
Data Residency
EU AI Act
Limited Access
GPT-5.6
By Falak Mahmood
Azure & Cloud
September 3, 2026

Claude Fable 5.1 in Foundry: cache math and the EU caveats

Claude Fable 5.1 keeps Fable 5's $10/$50 pricing but cuts cache reads to $0.25 per million tokens, which turns a 60-turn agent session from $17.25 into $10.50 and shrinks the premium over Opus 5 from 2x to about 22%. On Microsoft Foundry it ships Anthropic-hosted only, with no EU data zone, no Batches API, a zero default quota on pay-as-you-go, mandatory 30-day retention until Enterprise Frontier Safeguards arrive, and three breaking changes for teams migrating from Fable 5.

Claude Fable 5.1
Anthropic
Microsoft Foundry
Azure
Prompt Caching
AI Agents
Data Residency
GDPR
EU AI Act
LLM Cost
By Falak Mahmood
AI & Cloud Infrastructure
April 30, 2026

AI Agent Cost Economics: Why 100x and How to Cut It

Agent loops cost 10x to 50x what a chatbot interaction costs; multi-agent systems add another order of magnitude. The cost compounding is structural, not a bug. The cost reduction is structural too. Decomposing where the tokens go and how to bring agent economics back from runaway to acceptable.

AI Agents
Cost Optimization
LLM Cost
Agentic AI
Tokens
By Falak Mahmood
AI & Cloud Infrastructure
January 20, 2026

Prompt Caching: Cutting LLM Costs Without Quality Loss

A technical guide to prompt caching across Claude, Azure OpenAI, and GPT — what belongs in the cache, how to structure cache breakpoints, TTL realities, hit-rate optimization, and the anti-patterns that erase the savings.

Prompt Caching
LLM Cost
Claude
OpenAI
Optimization
By Falak Mahmood