The number that reset expectations

On January 20, 2025, DeepSeek released R1 under an MIT licence with API pricing of $0.55 per million input tokens on a cache miss, $0.14 on a cache hit, and $2.19 per million output tokens, and claimed performance on par with OpenAI o1. Whether the parity claim held on every benchmark was argued about for weeks. What nobody argued about was the price. A reasoning model that produced long chains of thought had been an expensive product, and now there was one you could run at a few dollars per million tokens and download the weights of.

The base model underneath it had already made the same point a month earlier. DeepSeek-V3 shipped on December 26, 2024, as a 671 billion parameter mixture of experts with 37 billion active per token, trained on 14.8 trillion tokens. After its promotional period ended on February 8 it settled at $0.27 per million input on a miss, $0.07 on a hit, and $1.10 output. Those are the numbers we want to trace forward, because they were the start of a curve rather than a one-off.

How the curve kept going

In September 2025 the V3.2-Exp release cut prices roughly in half again. VentureBeat reported the new rates as $0.028 per million on a cache hit, down from $0.07, $0.28 on a miss, down from $0.56, and $0.42 output, down from $1.68. The mechanism was architectural. DeepSeek Sparse Attention uses what the company called a lightning indexer to select only the most relevant tokens for attention, which flattens the cost curve as context grows toward the 128,000-token limit. Benchmarks held roughly steady against the previous version, with MMLU-Pro at 85.0, Codeforces up to 2121, and GPQA-Diamond slightly down at 79.9.

Then on April 24, 2026, V4 arrived in two sizes. V4 Flash is 284 billion total parameters with 13 billion active, a million-token context, and pricing of $0.14 input and $0.28 output per million, with cache hits at $0.0028 after the April 26 cut to one tenth of standard input. V4 Pro is 1.6 trillion total with 49 billion active at $1.74 and $3.48. The older R1 and V3.2 endpoints are being folded into these tiers, with the legacy aliases retiring on July 24.

Put those side by side and the shape is clear. Output on the fast tier went from $1.10 at V3 to $0.42 at V3.2 to $0.28 at V4 Flash, while the context window went from 128,000 to a million tokens. Cached input went from $0.07 to $0.0028, a 25x drop, and that last figure is the one that matters most for agents, which resend the same long prefix on every turn.

Bar chart of DeepSeek API pricing per million tokens, log scale, across four releases. V3 (December 2024): $0.27 input miss, $0.07 input hit, $1.10 output. R1 (January 2025): $0.55, $0.14, $2.19. V3.2-Exp (September 2025): $0.28, $0.028, $0.42. V4 Flash (April 2026): $0.14, $0.0028, $0.28.
DeepSeek API pricing per million tokens across four releases. R1 priced above V3 on launch; the decline resumed with V3.2-Exp and V4 Flash.

Three mechanisms, not one

It is tempting to tell this as a story about one lab undercutting the market. The pricing moves came from three separate engineering choices, and each one is available to anybody. The first is mixture of experts. V3 activates 37 of 671 billion parameters per token and V4 Flash activates 13 of 284 billion. Serving cost scales with active parameters, so a sparse model buys most of the quality of a dense one at a fraction of the compute per token.

The second is caching. Every DeepSeek price list since V3 has distinguished cache hits from misses, and the gap has widened from 4x to 50x. An agent loop that keeps a stable system prompt, tool definitions and conversation history pays the miss price once and the hit price thereafter. Teams that restructure prompts so the stable part comes first see their bill fall by an order of magnitude with no model change at all. The third is routing. Once a Flash-class model sits at $0.28 output, the economical design sends every request there first and escalates to a Pro-class model only when a cheap check says the answer was not good enough.

For comparison, CloudZero lists current list prices for the frontier tier at $2.50 and $15.00 for GPT-5.4, $3.00 and $15.00 for Claude Sonnet 4.6, and $5.00 and $25.00 for Claude Opus 4.7. Those models are stronger on the tasks we care about, and we use them. But the gap between them and the cheap tier is now 50x on output, and that gap is what a routing layer monetises.

What stops being the constraint

When we started costing applied projects in 2023, model spend was usually the first line on the budget and often the reason a project was declined. In 2026 it rarely is. A pipeline that classifies a million support tickets a month at a few hundred tokens each costs tens of dollars on the Flash tier. The money and the time now go to three other things.

Latency is the first. A million-token context at any price is useless for an interactive product if the first token arrives in thirty seconds, and the cheap tiers are not always the fast tiers. Integration is the second. Getting the model access to the right data, with the right permissions, and getting its output back into a system that trusts it, takes the engineering hours that model cost used to take. Evaluation is the third and the one that worries us. When a call costs a fraction of a cent it is very easy to make ten times as many of them, and very easy to skip measuring whether the answers are right, because nothing on the invoice is telling you to look.

Build, buy, and run it yourself

The open weights change the build-versus-buy question in a way the API prices alone would not. With MIT-licensed weights from R1 onward, and V4 Flash small enough in active parameters to serve on a modest cluster, teams with a compliance reason to keep data on their own hardware can do so at a per-token cost that would have been unthinkable in 2024. Whether they should is a separate question. Running a serving stack is a job, and the API price is now so low that the hardware amortisation rarely wins unless data residency forces the decision.

The on-device version of the same question is more interesting and we do not have a settled view. The distilled R1 variants at 32B and 70B were described by DeepSeek as on par with o1-mini, and a year of further distillation has pushed that frontier down. Our guess is that within another year the routing layer we described above will have a local model as its first tier, and the cloud will only see the escalations. We would like to see someone publish the accuracy loss curve for that arrangement on a real task, because the cost curve is no longer the interesting one.

Sources

  1. DeepSeek: DeepSeek-R1 release (January 20, 2025)
  2. DeepSeek: DeepSeek-V3 release (December 26, 2024)
  3. VentureBeat: DeepSeek V3.2-Exp cuts API pricing in half
  4. CloudZero: DeepSeek pricing (2026)