Inside Kimi K3: A 2.8-Trillion-Parameter Model That Uses Barely 2% of Itself
Moonshot AI released Kimi K3 on July 16-17, 2026 as a 2.8 trillion parameter Mixture-of-Experts model the company bills as the world's first open 3T-class system [1]. The model's efficiency, not just its size, is the real engineering story: K3 activates only 16 of its 896 experts per token, roughly 1.8% of its total parameters, while offering a 1 million token context window and native vision [1]. Two new architectural components, Kimi Delta Attention and Attention Residuals, underpin that sparsity [1]. The payoff showed up immediately on independent leaderboards: K3 debuted at number one on Arena.ai's Frontend Code Arena with 1,679 points, ahead of Claude Fable 5's 1,631 and GPT-5.6 Sol's 1,618 [2]. Moonshot itself is more measured about the bigger picture, acknowledging that K3's overall performance still trails Claude Fable 5 and GPT-5.6 Sol even as it claims frontier-level results on coding and agentic benchmarks specifically [1]. Pricing reinforces the positioning as a value play rather than an outright leader: $0.30 per million cache-hit input tokens, $3 per million cache-miss input tokens, and $15 per million output tokens, more expensive than DeepSeek V4 and GLM-5.2 but roughly a third of Claude Fable 5's cost [1]. The hosted product went live immediately across Kimi.com, Kimi Work, Kimi Code, and the Kimi API, with full model weights scheduled for public release on July 27, 2026 [1].


