A retrain, not a rebuild: how DeepSeek bought agentic gains without a new model
DeepSeek-V4-Flash-0731 is not a new model release - it's a re-post-training job on top of the exact same 284B-parameter, 13B-active MoE architecture and 1M-token context window that shipped with the April 2026 preview [1]. The company's own changelog is explicit that architecture and size are identical, and only the post-training stack changed [1]. Yet the benchmark movement produced by that retrain is large enough to look like a generational jump: Terminal Bench 2.1 climbed from 61.8 to 82.7, and DeepSWE - a harder agentic coding benchmark - jumped from 7.3 to 54.4 [2].
OfficeChai's analysis frames this precisely as a training-only upgrade rather than a new model [3], and BiGGo's reporting traces the mechanism to reinforcement learning and distillation-based tuning applied to the existing weights rather than a fresh pretrain [4]. That distinction matters: DeepSeek is demonstrating it can buy large agentic-capability gains through post-training alone, without the capital cost of training a new base model - a cheaper, faster lever than rivals who ship new architectures to chase the same gains.



