Business & Strategy

DeepSeek's 4x price rise: rethinking cheap open models

Av Technspire TeamAugust 17, 202614 visningar

On 13 August 2026, DeepSeek released V4-Pro (DeepSeek-V4-Pro-0813) to general availability across its app, web interface and API, positioning it squarely at agent workloads with flexible reasoning effort settings and native support for the OpenAI Responses API. Three days later, at 16:00 UTC on 16 August, the other shoe dropped: new API pricing took effect. V4-Pro output tokens now cost $3.96 per million during peak hours, up from a flat $0.87. That is a 4.5x increase on the number most cost models are built around, and on the rarely-discussed cached-input rate the multiple is far steeper. Quartz put the headline increase at up to 1,100 percent.

If you run an Azure-first stack in Sweden or elsewhere in the EU, this matters even if you never call DeepSeek's API directly. DeepSeek's promotional pricing has been the anchor for the entire "cheap open model" tier since early 2025. Budget spreadsheets, build-versus-buy decisions and vendor negotiations all referenced it. Those references are now stale, and the new peak window happens to cover the Swedish working morning. Time to redo the math, and while the spreadsheet is open, to revisit the data governance questions that the old prices made easy to defer.

What actually changed on 13 and 16 August

The model: V4-Pro goes GA with an agent focus

DeepSeek's release notes describe major agent upgrades with strong production gains. V4-Pro and V4-Flash gain a flexible reasoning-effort control: low for simple tasks, high for daily agent workflows, max for complex problems. The API adds native OpenAI Responses API support, which lowers the switching cost for teams already coded against that interface. In the consumer app and web product, V4-Pro surfaces as an "Expert Mode"; API model names are unchanged.

Third-party model trackers list V4-Pro-0813 at a 1M-token context window with outputs up to 384K tokens, and a Terminal Bench 2.1 score of 87.9, a strong result for terminal-driven agent tasks and the reason the release is being read as an agent play rather than a chat play. Long agent trajectories with large outputs are exactly the workloads the new pricing hits hardest, which is unlikely to be a coincidence.

The prices: peak and off-peak, effective 16 August

The new scheme splits the week into peak and off-peak periods. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday through Friday. Everything else is off-peak, billed at half the peak rate. DeepSeek's stated rationale: peak and off-peak pricing lets it "allocate resources more reasonably." The published rates per million tokens:

  • V4-Pro output: was $0.87 flat; now $3.96 peak, $1.98 off-peak.
  • V4-Pro input (cache miss): was $0.435; now $1.32 peak, $0.66 off-peak.
  • V4-Pro input (cache hit): was $0.003625; now $0.044 peak, $0.022 off-peak. That is roughly a 12x increase at peak, the source of the 1,100 percent headlines.
  • V4-Flash output: was $0.28; now $1.32 peak, $0.66 off-peak.
  • V4-Flash input: cache miss now $0.44 peak, $0.22 off-peak; cache hit $0.014 peak, $0.007 off-peak.

Some history makes the reversal sharper. As Engadget notes, DeepSeek's discounted rates were originally promotional and due to expire on 31 May; the company then announced the discounts would become permanent, and has now reversed course again. Any procurement narrative that treated DeepSeek pricing as a stable floor should be retired.

The arithmetic for a Swedish workday

The peak window is defined in UTC, which maps awkwardly for European teams. During Swedish summer time (CEST, UTC+2), peak hours run 03:00 to 06:00 and 08:00 to 12:00 local time on weekdays. In other words, interactive agent traffic generated by your users between morning standup and lunch is billed at the full peak rate. Your afternoon is off-peak. The window aligns with the Chinese business day, and European mornings simply happen to overlap it.

A worked example, ours rather than anyone's benchmark. Take an internal agent platform that generates 10 million V4-Pro output tokens per working day, spread evenly across an 08:00 to 17:00 CEST day. Four of those nine hours are peak. Before 16 August the output bill was 10M x $0.87, about $8.70 per day regardless of clock time. Now roughly 4.4M tokens land at $3.96 and 5.6M at $1.98, around $28.50 per day, more than a threefold increase without any change in usage. Add the input side: agent loops are heavy on cached context reads, and the cache-hit rate just went up by an order of magnitude. Teams that engineered aggressively around DeepSeek's near-free cache hits, with long shared system prompts and tool schemas, see the biggest percentage jump of all.

Perspective still matters. Engadget's comparison puts OpenAI's GPT-5.6 Sol at $30 per million output tokens and Moonshot's Kimi K3 at $15, so V4-Pro at $3.96 peak remains several times cheaper than Western frontier pricing. Only OpenAI's budget GPT-5.6 Luna, at $1.20, undercuts V4-Flash's new $1.32 peak output rate. DeepSeek has not become expensive. It has become normal, and normal changes the decision calculus for EU buyers, because the extreme discount was carrying costs of a different kind.

Why the ultra-cheap era was always ending

Three structural reads, clearly labelled as our analysis. First, peak and off-peak billing is what a capacity-constrained provider does. Uniform low prices fill your GPUs with price-insensitive batch traffic exactly when interactive demand spikes; time-of-day pricing pushes the batch work to the night shift. Second, the agent era inverts the token economics that made loss-leader pricing tolerable. Chat workloads are input-heavy and output-light; agent workloads with thinking modes and 384K-token outputs are the opposite, and output is the expensive direction. A model marketed on Terminal Bench results invites precisely the traffic that costs the most to serve. Third, promotional pricing did its job. DeepSeek's discounts reset global price expectations and forced responses from every competitor during 2025. Once the market position is established, the discount is pure cost.

The planning consequence: treat any headline-grabbing low price from any provider as a promotional artifact with a half-life, not a parameter you can build a three-year TCO model on. Model API pricing has repriced repeatedly in both directions since 2023. Your architecture should assume prices move; portability is the hedge, and it is cheaper to build in from the start than to retrofit under budget pressure.

Decision framework: stay, shift, or switch

Before choosing, quantify four things: what share of your model spend goes to DeepSeek endpoints; what fraction of that traffic falls in the 01:00-04:00 and 06:00-10:00 UTC weekday window; how dependent your unit economics are on cache-hit pricing; and whether the workload touches personal data or confidential business data at all. The fourth question dominates the other three for most Swedish enterprises.

  • Stay and optimise if the workload is batch-friendly and data-benign. Move scheduled jobs, evaluations and bulk processing entirely into off-peak hours, which is straightforward from a Swedish afternoon onwards. Drop reasoning effort to low where the task allows; the new control exists precisely so you can pay for thinking only where it earns its keep. Expect roughly half the peak rate for disciplined off-peak usage.
  • Shift hosting, keep the model family if you like the weights but not the API terms. DeepSeek has published open weights for earlier flagship models, and Microsoft has offered DeepSeek models (R1 onwards) through Azure AI Foundry, hosted in Microsoft's infrastructure, since early 2025. Foundry-hosted or AKS-self-hosted weights decouple you from both DeepSeek's price changes and its data-handling terms. Check availability and region support for the V4 generation before committing; GPU capacity for a large mixture-of-experts model is its own cost line and needs honest sizing.
  • Switch models if DeepSeek was chosen on price alone. At $3.96 peak output, the gap to alternatives you can buy with EU data residency, an Azure enterprise agreement and familiar compliance paperwork has narrowed substantially. Rerun your evaluation suite against the current Azure AI Foundry catalogue at current prices before assuming the 2025-era ranking still holds.
  • In every case, add price mobility. Route model calls through an abstraction you control (a gateway or a thin internal SDK), keep evaluation suites runnable against multiple providers, and put a quarterly reprice check in the platform team's calendar. The Responses API compatibility DeepSeek just shipped cuts switching costs in both directions; use that fact in your next negotiation.

The Swedish and EU angle: the discount was doing compliance work too

For EU enterprises, DeepSeek's hosted API has always carried governance questions that the pricing made tempting to postpone. The service processes prompts and outputs on infrastructure in China, a jurisdiction without an EU adequacy decision, which puts any personal data in those prompts into GDPR Chapter V transfer territory with weak practical mitigations. Italy's data protection authority ordered the DeepSeek app blocked in early 2025, and several other European regulators opened inquiries. None of that was resolved by low prices; it was outweighed by them, at least in informal engineering decisions that never reached a DPO's desk. At 4.5x the price, the trade reads differently. If a workload was only defensible because it was nearly free, it was never really defensible.

The cleaner EU pattern has been available for a while: run open weights on infrastructure you or Microsoft control. Weights are software; they carry no transfer question, no foreign-jurisdiction API terms and no exposure to a provider's pricing announcements. Azure AI Foundry hosting keeps inference inside Microsoft's datacentres under your existing Azure agreement, with the usual region selection and, for eligible services, EU Data Boundary commitments. Self-hosting on AKS with GPU node pools goes further and costs more in engineering. Under the AI Act, your deployer obligations do not disappear either way, but your documentation gets simpler when the data flow diagram stops at an EU region.

Procurement teams should also log the meta-lesson. A supplier that announced permanent discounts and then withdrew them within months is a supplier whose pricing is a strategy variable, not a commitment. For upphandling and vendor risk assessments, that argues for shorter pricing assumptions, explicit repricing clauses where you have contractual leverage, and architectural exit options where you do not. The peak window sitting on Chinese business hours is a reminder of whose demand curve sets the price; European buyers are along for the ride.

Takeaways

  • Reprice now. V4-Pro output went from $0.87 flat to $3.96 peak and $1.98 off-peak on 16 August; cache-hit input rose roughly 12x at peak. Update every cost model that references DeepSeek rates.
  • Map your traffic to the peak window. Peak is 01:00-04:00 and 06:00-10:00 UTC on weekdays, which covers Swedish mornings. Move batch work to off-peak for an immediate 50 percent saving on the new rates.
  • Use the reasoning-effort dial. Low effort for simple tasks is now a cost control, not just a latency control.
  • Reopen the data governance file. China-hosted inference on personal or confidential data was a transfer problem at $0.87 and remains one at $3.96. Decide deliberately this time.
  • Price the Azure-hosted alternative. Compare Foundry-hosted open weights and current catalogue models against DeepSeek's new rates with your own evaluation suite, not 2025 rankings.
  • Build price mobility into the platform. Gateway abstraction, portable evaluations, quarterly reprice reviews. The next surprise announcement should cost you a config change, not a quarter.

Sources