Inside the Architecture: How KDA and MLA Deliver 2.5x Efficiency
Kimi K3's headline number isn't just its 2.8 trillion total parameters - it's how few of them fire on any given token. The model activates 104 billion parameters through a mixture of 896 experts with only 16 active per token, spread across 93 layers [1]. Rather than stacking more standard attention layers, Moonshot built a hybrid stack of 69 Kimi Delta Attention (KDA) linear-attention layers and 24 Gated MLA layers, a design SGLang and the Miles RL training framework backed with simultaneous day-zero support [2]. On SGLang's disaggregated serving setup, the model reached 423 tokens per second at batch-1 decode with speculative decoding and 2,808 tokens per second per GPU [2]. The efficiency case matters because it's also winning benchmarks outright - Kimi K3 took the #1 spot on the Frontend Code Arena at 1,679 Elo, ahead of Claude Fable 5 and GPT-5.6 Sol, and posted the best open-weight GPQA Diamond score ever published [1][3]. It still trails the very top closed models on Artificial Analysis's broader Intelligence Index, landing 4th overall - a reminder that 'best open-weight' and 'best overall' aren't yet the same claim [3].



