An Upgrade Built Entirely in Post-Training
DeepSeek's most consequential decision with V4-Flash-0731 was what it didn't touch: the architecture. The model keeps the exact same 284 billion total parameters, 13 billion active in its mixture-of-experts design, and the same 1 million-token context window as the April preview [2]. Every benchmark gain came from a heavier round of post-training focused specifically on coding, agentic workflows, and tool use, rather than any new pretraining run [1].
The jumps are large enough to look like a new model class. Terminal-Bench 2.1 climbed from 72.1 (on the prior V4-Pro preview) to 82.7, NL2Repo rose from 38.5 to 54.2, Cybergym from 52.7 to 76.7, and Toolathlon-Verified from 55.9 to 70.3 [1]. That's a striking result for anyone who assumes frontier progress requires scaling up parameters or compute - here, the same weights, retrained differently, closed much of the gap to larger and more expensive systems.


