DeepSeek V4's $0.28 Agentic Output Floor: What Heavy AI Users Pay Now
DeepSeek V4 set a $0.28 per million output token floor for agentic workloads, undercutting OpenAI Luna by 77%. Here is how this resets your agent cost math.
On July 31, 2026, DeepSeek did not just launch another cheap model. It moved the goalposts of the entire AI price war by pricing agentic output at $0.28 per million tokens. That single number is a structural ceiling for the whole industry, and it lands two days after OpenAI cut GPT-5.6 Luna prices by 80%. For heavy AI users paying $300 or more a month, this is the moment the economics of agentic automation stopped being a premium and became a commodity.
The release is officially DeepSeek-V4-Flash-0731, a re-post-trained version of the V4 Flash line that was already the cheapest credible coding agent on the market. Input stays at $0.14 per million tokens and output at $0.28 per million tokens. Nothing changed on price. What changed is that DeepSeek is now openly using that output price as a strategic weapon against every US proprietary model that still charges flagship rates for agent work.
Output, not input, is now the battleground
The old way to compare AI costs was input price. Prompts are long, everyone assumed input tokens would dominate the bill. Agentic workloads flipped that assumption. An agent sends tool calls, writes code, and streams reasoning, and the output side of the ledger explodes. When a coding agent runs for an hour, output tokens often outnumber input tokens by a wide margin.
DeepSeek is betting the market has caught up to this reality. By holding output at $0.28 per million tokens, it makes the expensive half of an agentic bill almost free. Against OpenAI’s GPT-5.6 Luna, priced at $0.20 in and $1.20 out, DeepSeek output is about 77% cheaper. Against Moonshot’s Kimi K3 at $3 in and $15 out, the gap is larger still. The cost of generating useful work per token has become the real price signal, and DeepSeek owns it.

The performance is no longer the excuse to pay more
The old counterargument was that cheap models deliver cheap results. DeepSeek V4 Flash dismantles that case on the agent benchmarks that matter. On Terminal Bench 2.1 it scores 82.7, on Cybergym 76.7, on DSBench-FullStack 68.7, and on DSBench-Hard 59.6. Its aggregate Intelligence Index sits around 50, against Luna’s 51. That is a gap so small it mostly disappears inside run-to-run variance.
For a heavy AI user, the practical read is simple. If a workload already passes on V4 Flash in your own task mix, the price-to-performance ratio is now absurdly good. The same agentic output that costs $1.20 per million on Luna costs $0.28 on DeepSeek. At the token volumes serious users run, that difference is thousands of dollars a month, and it compounds on every task the agent automates.
The one workload to keep off V4 Flash is work that needs the larger reasoning surface of a flagship model, where fewer output tokens buy more correct output. That is a capability question, not a cost question, and it is the only scenario where the premium survives on merit.

What this does to the rest of the market
DeepSeek positioning output at $0.28 as a hard ceiling forces Western labs into an uncomfortable trade. They can compress margins to match, or they can concede the high-volume, cost-sensitive agent segment. OpenAI’s 80% Luna cut on July 30, announced before the V4 Flash launch, already looked like a response to DeepSeek pressure. Two days later, DeepSeek undercut that cut anyway. That sequencing is the clearest evidence yet that the price war has moved to agentic output pricing and that nobody is setting the floor except the cheapest credible player.
For developers and FinOps teams, the near-term wrinkle is that full V4 is still coming. Sources point to the official V4 release on or around August 3, which may bring its own pricing structure and possibly higher rates for the larger model. If you are planning around V4 Flash pricing, budget the flash tier now and treat the flagship tier as a separate decision once official numbers land. V4-Pro’s Responses API support is already slated for early August, so the transition window is short.
The bottom line for heavy AI users
The deepseek v4 pricing story is not a discount. It is a redefinition of what agentic work should cost. When output tokens become a utility line item instead of a premium, the users who win are the ones who route high-volume agent workloads to the cheap output pool and reserve the pricey models for the tasks that genuinely need them.
Move your batchable agentic tasks to DeepSeek V4 Flash and the per-task cost collapses to a level no US provider matches today. Benchmark it against your own prompt mix first, because benchmark spreads of a point or two on the leaderboard do not translate into your workflow. But for most heavy users, the arithmetic is no longer close. The agentic output ceiling is set, and it is emerald green.
If you are already tracking spend across providers, this is the week to reweight your routing rules. Separate the workloads that can run on a $0.28 output model from the ones that cannot, and let the floor do the cost cutting for you.
Now available
Stop guessing your AI limits
The Mac app and web dashboard watch your Claude, ChatGPT, Gemini and more, and warn you before quotas hit.