The Architecture Trick: 763 Billion Parameters, Only 8-16 Billion Active
V4.1-Flash's headline number is deceptive. The model carries a 552B-parameter backbone plus 196B additional Engram parameters for a 763B total[1], but a new Causal-Encoder-Decoder design - a 40-layer Transformer split into a 20-layer causal encoder and a 20-layer decoder - projects the decoder's global KV cache from the encoder's final hidden states instead of recomputing it at every decoder layer[1]. That lets the model activate just 8B parameters per token during prefill and 16B during decode[1], regardless of how large the underlying weights are. Paired with FP4 KV caching (E2M1 format) and Compressed Sparse Attention 2, the global KV cache footprint drops to 890 bytes per token - about a quarter of DeepSeek-V4-Flash's footprint and roughly 437x smaller than DeepSeek-V1's[2]. The context window stretches to 1 million tokens, and DeepSeek says scaling from 4K to 1M tokens adds only about 25% extra computation[3]- a sign the architecture, not just brute-force compute, is doing the work.



