DeepSeek V4.1-Flash launch
TECH

DeepSeek V4.1-Flash launch

45+
Signals

Strategic Overview

  • 01.
    DeepSeek officially released DeepSeek-V4.1-Flash on September 10, 2026, a multimodal Mixture-of-Experts model with a 552B-parameter backbone plus 196B additional Engram parameters (763B total) and a 1M-token context window, built on a new Causal Encoder-Decoder architecture (a 40-layer Transformer split into a 20-layer causal encoder and a 20-layer decoder).
  • 02.
    The design lets the model activate only 8B parameters per token during prefill (reading input) and 16B during decode (generating output), regardless of the model's full size, by projecting the decoder's global KV cache from the encoder's final hidden states.
  • 03.
    Starting September 14, 2026 at 04:00 UTC, all requests to deepseek-v4-pro will be automatically routed to V4.1-Flash and billed at Flash-series rates, effectively retiring V4-Pro; legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are also temporarily routed to V4.1-Flash.
  • 04.
    Model weights are open on Hugging Face under the MIT License with no access restrictions, supporting deployment via vLLM, SGLang, and Transformers.
  • 05.
    V4.1-Flash beats or matches GPT-5.6 Sol and Claude Opus-5 on several agentic and coding benchmarks, scoring 90.6 on Terminal-Bench 2.1 (vs 88.8 and 89.1) and 74.2 on DeepSWE v1.1 (vs 73.0 and 74.0).
  • 06.
    Off-peak API pricing was cut to RMB 0.02 per million tokens for cache-hit input, RMB 1 for cache-miss input, and RMB 4 for output, with peak-hour pricing at double the off-peak rate.

The Architecture Trick: 763 Billion Parameters, Only 8-16 Billion Active

V4.1-Flash's headline number is deceptive. The model carries a 552B-parameter backbone plus 196B additional Engram parameters for a 763B total[1], but a new Causal-Encoder-Decoder design - a 40-layer Transformer split into a 20-layer causal encoder and a 20-layer decoder - projects the decoder's global KV cache from the encoder's final hidden states instead of recomputing it at every decoder layer[1]. That lets the model activate just 8B parameters per token during prefill and 16B during decode[1], regardless of how large the underlying weights are. Paired with FP4 KV caching (E2M1 format) and Compressed Sparse Attention 2, the global KV cache footprint drops to 890 bytes per token - about a quarter of DeepSeek-V4-Flash's footprint and roughly 437x smaller than DeepSeek-V1's[2]. The context window stretches to 1 million tokens, and DeepSeek says scaling from 4K to 1M tokens adds only about 25% extra computation[3]- a sign the architecture, not just brute-force compute, is doing the work.

Beating GPT-5.6 Sol and Claude Opus 5 at a Fraction of the Cost

Beating GPT-5.6 Sol and Claude Opus 5 at a Fraction of the Cost
DeepSeek V4.1-Flash edges out GPT-5.6 Sol and Claude Opus 5 on Terminal-Bench 2.1 and DeepSWE v1.1.

DeepSeek is making explicit benchmark claims against the two reigning closed-model leaders. On Terminal-Bench 2.1, V4.1-Flash scores 90.6 versus 88.8 for GPT-5.6 Sol and 89.1 for Claude Opus-5; on DeepSWE v1.1 it scores 74.2 versus 73.0 and 74.0 respectively[2]. Those are narrow margins, but they arrive alongside off-peak API pricing of RMB 0.02 per million cache-hit input tokens, RMB 1 for cache-miss input, and RMB 4 for output (peak-hour rates double)[4]- a fraction of what frontier-lab pricing typically runs. The model isn't only chasing OpenAI and Anthropic either; coverage names Moonshot AI's Kimi K3 as the direct rival in the open, low-cost model segment[5], underscoring that the sharpest competition right now is happening among fast-moving open-weight players as much as against the closed labs. Independent testing has picked up on this quickly - one design-focused benchmarker found V4.1-Flash reaching 98% of a leading closed model's score at roughly 1.4% of the cost on everyday design tasks, framing it as evidence that open models are closing the gap with closed ones.

Retiring V4-Pro Is Not Optional

This isn't a parallel release that leaves the old model running. Starting at 04:00 UTC on September 14, 2026, every request to deepseek-v4-pro will be automatically routed to V4.1-Flash and billed at Flash-series rates, effectively retiring V4-Pro; the legacy deepseek-v4-flash and deepseek-v4-flash-vision-exp names are also temporarily routed to the new model[6]. DeepSeek says its own internal and external testing shows V4.1-Flash surpassing V4-Pro on performance, cost, speed and total completion time[4], which is the stated justification for forcing the migration rather than offering it as an option. In practice, every V4-Pro integration gets switched to a materially different architecture within days of announcement, whether or not a given workload has been validated against it.

Compared To the Last DeepSeek Shock, This One Barely Registered

The context that makes this release notable isn't purely technical. In January 2025, DeepSeek's V3 and R1 release triggered a historic single-day selloff, with Nvidia's share price falling about 17% - roughly a $600 billion market-cap loss, the largest one-day drop for a single company in US stock market history at the time[7]. V4.1-Flash's launch, matching or beating two current frontier closed models on coding benchmarks at a fraction of the price, produced no comparable market disruption[7]. That gap is itself notable: eighteen months after the original shock, a Chinese lab shipping open, MIT-licensed weights that compete with top closed models has become an expected cadence rather than a surprise, following DeepSeek's pattern of splitting its main model line into separate Pro and Flash variants earlier in the V4 generation[8].

The Catch: You Still Need a Server Farm to Run It Yourself

The MIT-licensed weights are freely downloadable from Hugging Face with support for vLLM, SGLang, and Transformers, and no access restrictions gate who can use them[1]. But openness doesn't mean accessibility: a 552B-backbone, 763B-total-parameter model still requires server-grade, multi-GPU hardware to self-host[1], even with the KV-cache optimizations that make serving it at scale cheaper. The efficiency gains mostly benefit whoever is already running the infrastructure - DeepSeek's own API, or well-resourced enterprises - rather than putting a frontier-class model within reach of a single workstation.

Historical Context

2023-11-02
DeepSeek released its first public model family, DeepSeek Coder.
2025-01
DeepSeek released V3 and the R1 reasoning model, triggering a market shock in which Nvidia's share price fell about 17% in a single day, the largest one-day market-cap loss for a single company in US stock market history at the time.
2026
DeepSeek's V4 generation split its main model line into separate Pro and Flash variants, with V4.1-Flash now released to replace and outperform V4-Pro.
2026-09-10
V4.1-Flash formally launched with open MIT-licensed weights on Hugging Face and new Flash-tier API pricing effective the same day.

Power Map

Key Players
Subject

DeepSeek V4.1-Flash launch

DE

DeepSeek

Developer and publisher of V4.1-Flash; positions it as the smallest model in a new CED architecture family designed to scale to larger models, retiring V4-Pro in favor of it and cutting Flash-tier API prices.

OP

OpenAI (GPT-5.6 Sol)

Referenced as the benchmark rival DeepSeek claims to beat or match on coding and agentic tasks at far lower per-token pricing, narrowing OpenAI's differentiation.

AN

Anthropic (Claude Opus 5)

Referenced as the other benchmark rival DeepSeek claims to match or edge out on coding benchmarks like DeepSWE v1.1 and Terminal-Bench 2.1.

MO

Moonshot AI (Kimi K3)

Direct Chinese competitor whose Kimi K3 model is named as a rival V4.1-Flash competes against in the open, low-cost model segment.

HU

Hugging Face

Distribution platform hosting the open MIT-licensed V4.1-Flash weights for self-hosted deployment.

DE

Developers and API users running agentic and coding workloads

Primary beneficiaries of the KV-cache reduction, since a lower memory footprint and higher throughput cut the cost of long-running, input-heavy agent jobs.

Fact Check

8 cited
  1. [1] DeepSeek-V4.1-Flash (Hugging Face model card)
  2. [2] DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache and Cross-Layer Attention Reuse
  3. [3] DeepSeek V4.1 Flash Explained: How It Cuts AI Memory 8x
  4. [4] DeepSeek Formally Launches V4.1-Flash, Routes V4-Pro Requests to Flash
  5. [5] DeepSeek V4.1 Flash Model Launch
  6. [6] DeepSeek API Updates
  7. [7] Why DeepSeek Didn't Cause an Investor Frenzy Again in 2025
  8. [8] DeepSeek Timeline: Release Dates

Source Articles

Top 5

THE SIGNAL.

Analysts

Frames V4.1-Flash as evidence that AI competition is shifting from raw benchmark scores toward the operational cost of running long, agentic jobs, since context-carrying cost dominates economics as agents work longer. As agents work longer, AI economics increasingly depend on the cost of carrying context through the job.

Grant Harvey
Lead Writer, The Neuron

Argues that per-token intelligence matters less than the total cost of completing a task correctly: 'The smartest model per token may matter less than the cost of getting the whole job done correctly.'

Grant Harvey
Lead Writer, The Neuron
The Crowd

🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6

@@deepseek_ai22940

We benchmarked DeepSeek V4.1 Flash by @deepseek_ai. It reached 98% of GPT-6 Astra's score at 1.4% of the cost on everyday design tasks based on user requests. Every model except Astra scored lower AND cost more. Are open models overtaking closed ones? Full results below ⇘️

@@OpenDesignHQ9341

DeepSeek has open-sourced some new code repositories, making it easier to deploy V4.1 Flash as well as subsequent open-source models: https://github.com/deepseek-ai/deepseek-recipe https://github.com/deepseek-ai/DeepSelect https://github.com/deepseek-ai/DeepJIT

@@tianyi883
Broadcast
Deepseek V4.1 Flash (Fully Tested): 200 TPS & Beats Astra!? (+New Architecture Overview)

Deepseek V4.1 Flash (Fully Tested): 200 TPS & Beats Astra!? (+New Architecture Overview)

DeepSeek V4.1 Flash Is INSANELY GOOD! Fast, Cheap, Powerful! (Fully Tested)

DeepSeek V4.1 Flash Is INSANELY GOOD! Fast, Cheap, Powerful! (Fully Tested)

🚨 DeepSeek V4.1 Flash VAZOU! 1M de Contexto, Visão e Velocidade ABSURDA!

🚨 DeepSeek V4.1 Flash VAZOU! 1M de Contexto, Visão e Velocidade ABSURDA!