Posts tagged with "LLM Pricing"

Found 7 posts

Business & Strategy
August 17, 2026

DeepSeek's 4x price rise: rethinking cheap open models

DeepSeek's V4-Pro reached general availability on 13 August 2026 with a 1M-token context and strong agent benchmarks, and three days later its peak-hour output price rose from a flat $0.87 to $3.96 per million tokens. For EU teams that built agent cost models around ultra-cheap open-weight APIs, the arithmetic, the data governance questions and the Azure hosting options all deserve a fresh look.

DeepSeek
LLM Pricing
TCO
Open Models
Azure AI Foundry
Data Residency
GDPR
AI Strategy
EU Enterprises
By Technspire Team
AI & Machine Learning
August 14, 2026

Workhorse shootout: Gemini 3.7 Flash, GPT-5.6 Luna, Sonnet 5

Google shipped gemini-3.7-flash as generally available on 13 August 2026 at an introductory $0.75/$3.75 per million tokens, two weeks after OpenAI cut GPT-5.6 Luna by 80 percent and days after Anthropic locked Claude Sonnet 5 at $2/$10 permanently. We compare the three workhorse models on list price, context, Azure availability and EU residency, and show why cost per completed task beats cost per token.

Gemini 3.7 Flash
GPT-5.6 Luna
Claude Sonnet 5
Model Comparison
LLM Pricing
Azure OpenAI
Microsoft Foundry
Enterprise AI
By Technspire Team
Business & Strategy
August 12, 2026

LLM cost planning autumn 2026: Sonnet 5 stays at $2/$10

Anthropic has cancelled the Claude Sonnet 5 price increase scheduled for 1 September 2026, making the introductory $2 input / $10 output per million tokens the permanent standard price. For teams running Claude on Azure through Microsoft Foundry, that removes a planned 50% jump from autumn budgets and reshapes the mid-tier price comparison against GPT-5.6 Terra and Gemini 3.1 Pro.

LLM Pricing
Claude Sonnet 5
Azure AI Foundry
Microsoft Foundry
GPT-5.6
Gemini
Cost Optimization
AI Budgeting
Anthropic
By Technspire Team
AI & Machine Learning
July 27, 2026

Claude Opus 5 for long-running agents: the cost math

Claude Opus 5 launched on 24 July 2026 at $5/$25 per million tokens with a 1M context window and day-one availability in Microsoft Foundry. For long-running agents the per-token price is the wrong unit: we work through cost per completed task against Sonnet 5 and GPT-5.6 Sol, and flag the EU data-residency caveat Swedish Azure teams need to check first.

Claude Opus 5
Anthropic
AI Agents
Microsoft Foundry
Azure
LLM Pricing
Claude Sonnet 5
GPT-5.6
Cost Optimization
Data Residency
By Technspire Team
AI & Machine Learning
July 23, 2026

Tokens are the new pricing lever: Gemini 3.6 Flash math

Google's 21 July release of Gemini 3.6 Flash pairs an output-price cut from $9.00 to $7.50 per million tokens with a claim of roughly 17% fewer output tokens per task, compounding to about 31% lower output cost for unchanged work. That combination makes per-million-token price sheets unreliable for model comparison, and Azure teams should measure cost per completed task instead.

Gemini
Google
LLM Pricing
Token Efficiency
Azure OpenAI
Azure AI Foundry
FinOps
Cost Optimization
AI Procurement
By Technspire Team
AI & Cloud Infrastructure
July 15, 2026

Running GPT-5.6 the enterprise way on Microsoft Foundry

GPT-5.6 (Sol, Terra, Luna) went GA in Microsoft Foundry on 9 July 2026, day-and-date with OpenAI, alongside a new Asia-Pacific Data Zone and a hosted agents runtime with VNet integration. A practical guide for Swedish and EU Azure teams: choosing between the three models, picking Global Standard versus EU Data Zone versus PTUs, worked cost math on the launch prices, and a two-week adoption checklist.

GPT-5.6
Microsoft Foundry
Azure OpenAI
Data Zones
EU Data Residency
Hosted Agents
VNet
AI Agents
LLM Pricing
By Technspire Team
AI & Machine Learning
May 29, 2026

Claude Opus 4.8: the effort dial, fast mode and token math

Claude Opus 4.8 arrives at unchanged pricing with an effort control on all plans, a fast mode at a third of the previous fast-inference cost, and a Messages API change that lets system entries sit inside the messages array so mid-task instruction updates no longer invalidate the prompt cache. Worked token math shows cache hit rate remains the biggest cost lever, and a four-question framework matches effort, speed and fan-out to each workload.

Claude Opus 4.8
Anthropic
AI Agents
LLM Pricing
Prompt Caching
Token Economics
Claude Code
Fast Mode
By Technspire Team