DeepSeek-V4-Flash-0731 public beta launch and its role in the AI price war
TECH

DeepSeek-V4-Flash-0731 public beta launch and its role in the AI price war

29+
Signals

Strategic Overview

  • 01.
    DeepSeek released the official public beta of DeepSeek-V4-Flash-0731 on July 31, 2026, superseding the April preview version of the model.
  • 02.
    The model keeps the identical architecture and parameter count as the April preview (284B total / 13B active parameters, MoE, 1M-token context) - all performance gains come from an improved post-training pipeline, not new architecture or pretraining.
  • 03.
    Model weights are released under the MIT License and hosted on Hugging Face, alongside a DSpark speculative-decoding variant.
  • 04.
    API pricing is held at $0.14 per 1M input tokens and $0.28 per 1M output tokens with a deep cache-hit discount, though DeepSeek has signaled plans to introduce peak-hour pricing later that would double billing for 7 hours a day during Beijing business hours.

An Upgrade Built Entirely in Post-Training

DeepSeek's most consequential decision with V4-Flash-0731 was what it didn't touch: the architecture. The model keeps the exact same 284 billion total parameters, 13 billion active in its mixture-of-experts design, and the same 1 million-token context window as the April preview [2]. Every benchmark gain came from a heavier round of post-training focused specifically on coding, agentic workflows, and tool use, rather than any new pretraining run [1].

The jumps are large enough to look like a new model class. Terminal-Bench 2.1 climbed from 72.1 (on the prior V4-Pro preview) to 82.7, NL2Repo rose from 38.5 to 54.2, Cybergym from 52.7 to 76.7, and Toolathlon-Verified from 55.9 to 70.3 [1]. That's a striking result for anyone who assumes frontier progress requires scaling up parameters or compute - here, the same weights, retrained differently, closed much of the gap to larger and more expensive systems.

Why Holding the Price Steady Was the Real Weapon

DeepSeek didn't even need to cut prices to escalate the price war - it just held them. API pricing stays at $0.14 per million input tokens and $0.28 per million output tokens, with a cache-hit discount that brings repeat-context input down to roughly $0.0028 per million tokens, near a 98 percent discount [3]. Against Claude Opus 4.8's list price of $5 per million input and $25 per million output tokens, that's routinely cited as a 97 to 99 percent gap on cost [4].

The more interesting comparison is against OpenAI, not Anthropic. Even after OpenAI's own steep price cut on GPT-5.6 Luna, DeepSeek's per-task cost still comes in roughly 60 percent lower, while trailing GPT-5.6 Luna by just a single point on the Artificial Analysis Intelligence Index [5]. In other words, DeepSeek is now within a rounding error of a frontier closed model's general intelligence score at a fraction of the cost.

One detail complicates the 'cheapest ever' framing: DeepSeek has signaled it plans to introduce peak-hour pricing later, doubling all billing for seven hours a day during Beijing business hours [3]. No effective date has been announced, but it's a reminder that today's headline price is a promotional floor, not a permanent one.

The Gap Between the Leaderboard and the Terminal

Not everyone is convinced the benchmark gains translate cleanly into real-world coding work. DeepSeek's own reported numbers were produced using its own upcoming proprietary harness running in a minimal-effort mode, which means results measured through other developer tools may look different - a caveat independent testers have flagged directly rather than taking the numbers at face value.

That skepticism shows up in hands-on community testing too: some developers report DeepSeek's real coding performance landing meaningfully behind Claude Opus 4.8 and Sonnet on practical tasks, even where the published benchmark suggests near parity. The tension is less about whether V4-Flash-0731 improved - it clearly did, jumping 10 points on the Artificial Analysis Intelligence Index in one release [5]- and more about how much of that gain shows up once a different harness, a different prompt, or a production codebase enters the picture.

Why This Doesn't Look Like the R1 Shock

DeepSeek's R1 release in January 2025 rattled Western AI markets because it arrived as a genuine surprise - a low-cost reasoning model nobody had priced in. V4-Flash-0731 is landing in a different environment. The market has already spent a year absorbing the idea that a Chinese lab can ship frontier-adjacent capability at a fraction of Western pricing; Moonshot AI's Kimi K3 was already ahead of DeepSeek on the open-weight leaderboard without triggering anything like the R1 disruption [6].

That context matters for reading this release correctly. It's a genuine, measurable step forward - and it arrives backed by fresh capital, with DeepSeek having raised roughly $7.4 billion from investors including Tencent and NetEase at a 350 billion yuan valuation, funding the company has earmarked for doubling headcount and pushing further into agentic AI [6]. But it's an incremental escalation of an already-known competitive dynamic, not a new shock to reprice around. Notably, the release also arrived slightly behind DeepSeek's own mid-July target, and without the companion V4-Pro update some had expected alongside it [6].

Who Actually Has to Respond

The immediate pressure lands on Anthropic and OpenAI, but not evenly. OpenAI already moved first, cutting GPT-5.6 Luna pricing steeply, and DeepSeek's release still undercuts that discounted price by roughly 60 percent per task [5]. Anthropic hasn't matched either move; its bet, reflected in how its own user base talks about the comparison, is that Claude's surrounding product - the coding harness and agent tooling around the model, not the raw per-token price - is what customers are actually paying for.

Moonshot AI is the wrinkle in the 'DeepSeek wins' narrative: Kimi K3 still leads the open-weight frontier by roughly 7 points on the Artificial Analysis Intelligence Index, and Moonshot is scaling further with added Nvidia GPU capacity [6]- a reminder that DeepSeek is competing for a share of the open-weight lead, not holding it outright.

The MIT license is its own kind of pressure valve. Anyone can pull the weights from Hugging Face and self-host, though the hardware bar is real - full precision needs something like a 4xGB300 node, while 3-bit quantization brings it down to roughly 110GB [2]. That's out of reach for hobbyists but well within range of a well-resourced startup or research lab, which quietly shifts some of the competitive pressure away from API pricing entirely and onto who controls the best inference stack.

Historical Context

2023-11-02
Released DeepSeek Coder, its first public model family.
2025-01-20
Released DeepSeek-R1, a reinforcement-learning-trained reasoning model that shocked Western AI markets over its cost-efficiency.
2026-04-24
Unveiled the V4 Preview generation (V4-Pro and V4-Flash), its first flagship release a year after the R1 shock.
2026-07-24
Retired the legacy deepseek-chat and deepseek-reasoner API aliases and their off-peak discount pricing for V3/R1.
2026-07-31
Launched DeepSeek-V4-Flash-0731 as the official public-beta release, superseding the April preview with re-post-trained agentic and coding gains at unchanged pricing.

Power Map

Key Players
Subject

DeepSeek-V4-Flash-0731 public beta launch and its role in the AI price war

DE

DeepSeek

Developed and released V4-Flash-0731; recently raised roughly $7.4 billion in a private funding round backed by Tencent and NetEase at a 350 billion yuan valuation, earmarked for doubling headcount and pushing agentic AI development.

AN

Anthropic (Claude Opus 4.8)

Premium rival whose pricing ($5/M input, $25/M output) is undercut by roughly 97-99 percent by V4-Flash's rates, making it the central contrast in the price-war narrative.

OP

OpenAI (GPT-5.6 Luna)

Competing frontier lab; V4-Flash-0731 trails GPT-5.6 Luna by just 1 point on the Artificial Analysis Intelligence Index but still undercuts it on cost even after OpenAI's own steep price cut.

MO

Moonshot AI (Kimi K3)

Open-weight rival lab whose Kimi K3 model still leads the open-weight frontier by about 7 Intelligence Index points over V4-Flash-0731, while scaling with additional Nvidia GPU capacity.

TE

Tencent Holdings Ltd. / NetEase Inc.

Investors backing DeepSeek's roughly $7.4 billion private funding round, funding headcount growth and agentic AI R&D.

AR

Artificial Analysis

Independent benchmarking firm whose Intelligence Index comparison is the primary neutral yardstick used to place V4-Flash-0731 against GPT-5.6 Luna and other rivals.

Fact Check

6 cited
  1. [1] DeepSeek Upgrades DeepSeek-V4-Flash-0731 With Major Agentic and Coding Gains
  2. [2] DeepSeek-V4-Flash-0731 Model Card
  3. [3] DeepSeek V4-Flash-0731 Pricing Explained
  4. [4] Claude Opus 4.8 vs DeepSeek V4-Flash: API Pricing and Performance Compared
  5. [5] DeepSeek V4 Flash 0731 Scores 50 on the Artificial Analysis Intelligence Index, 10 Points Above Previous DeepSeek V4 Flash
  6. [6] DeepSeek Releases Official V4-Flash Model as China's AI Race Intensifies

Source Articles

Top 5

THE SIGNAL.

Analysts

Argues the model punches above its weight class relative to larger competitors and may currently be one of the best value-per-intelligence options available, based on his own testing via OpenRouter.

Simon Willison
Independent AI commentator and developer

Notes V4-Flash-0731 shares identical architecture and pricing with the earlier preview yet lands on the Pareto frontier for intelligence versus cost, remaining competitive with GPT-5.6 Luna even after OpenAI's own steep price cut.

Artificial Analysis
Independent LLM benchmarking organization
The Crowd

DeepSeek-V4-Flash Official API is now LIVE in public beta! We've massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! The official V4-Flash now natively supports the Responses API format and is fully adapted for Codex! Check out the configuration details in our official API docs: https://api-docs.deepseek.com/quick_start/agent_integrations/codex

@@deepseek_ai27042

DeepSeek V4 Flash just tied Gemini 3.6 Flash on the intelligence index. same score. 30x cheaper on output. Google charges $7.50 per million output. DeepSeek charges $0.28. and it's open weights. closed pricing makes no sense anymore

@@shiri_shh2620

The AI price war is ON. > DeepSeek just dropped V4 Flash 0731 (big capability jump, $0.28/M output) > OpenAI cut GPT-5.6 Luna by 80% to $1.20/M (OpenRouter is currently offering 50% off, bringing it to $0.60/M) Same budget tier now, so I ran both through 3 canvas tests: Rubik's cube: DeepSeek nails every rotation. Luna's turns are buggy, stickers fly off mid-move. Fireworks: both fine, but Luna lags hard when the burst explodes. Pen writing "Hello": neither is legible. DeepSeek's cursive is looser, but Luna gave up on animating, it rendered a static image. GPT-5.6 is great at coding — but IMO only on Sol. In the same pricing tier, DeepSeek clearly beats Luna.

@@stevibe1933

DeepSeek V4-Flash is officially out, still dirt cheap. USA don't like that, and want ban open source models.

@u/bi4key1800
Broadcast
DeepSeek V4 Flash Is OUT, OpenAI "mewthree" + GPT-5.6 Price/Speed Update, Qwen 3.8 Kinsley, & More!

DeepSeek V4 Flash Is OUT, OpenAI "mewthree" + GPT-5.6 Price/Speed Update, Qwen 3.8 Kinsley, & More!

Deepseek V4 Flash (0731 - Fully Tested): TOP 5 in my TESTS! This is AN ACTUAL COMEBACK!!!

Deepseek V4 Flash (0731 - Fully Tested): TOP 5 in my TESTS! This is AN ACTUAL COMEBACK!!!

DeepSeek V4 Flash 0731 Just DROP and Is a HUGE Win for Local AI (2 DGX Sparks!)

DeepSeek V4 Flash 0731 Just DROP and Is a HUGE Win for Local AI (2 DGX Sparks!)

DeepSeek-V4-Flash-0731 public beta launch and its role in the AI price war — AI News | Agentic Brew