Architecture Behind the Efficiency Claim
Kimi K3's headline number is 2.8 trillion total parameters, but the more consequential number is how few of them fire on any given token. The model activates just 16 of 896 experts per request - roughly 1.8% - under a new Stable LatentMoE framework, paired with two new mechanisms Moonshot calls Kimi Delta Attention (KDA) and Attention Residuals (AttnRes) [1]. Moonshot's own technical report frames the result as approximately a 2.5x improvement in scaling efficiency over Kimi K2, and describes K3 as trailing the very top proprietary models (Claude Fable 5, GPT-5.6 Sol) while still consistently outperforming other tested models across its evaluation suite [1]. The practical footprint matters too: quantized weights via MXFP4/MXFP8 shrink storage from 5.6 terabytes to about 1.56 terabytes, a roughly 72% reduction that makes self-hosting materially more feasible for organizations with serious infrastructure [2]. The active-parameter count per request lands around 104 billion out of 2.8 trillion total [2]- the story is sparsity, not brute-force scale.



