Efficiency by Necessity: The Architecture Behind K3's Leap
Moonshot AI's president Yutong Zhang frames Kimi K3's build strategy as a constraint turned into a method: "We knew we didn't have the luxury to simply scale up compute. That forced us to focus on fundamental research and efficiency." [1]That constraint produced K3's new architecture - Kimi Delta Attention, Attention Residuals, and Stable LatentMoE - stacked on a 2.8-trillion-parameter mixture-of-experts design [2], of which 104 billion parameters and only 16 of 896 experts activate per token [3]. Moonshot describes the result as "2.5x the intelligence per unit of compute, not just more params" [3]. A hands-on YouTube review of the release found that the fraction of experts firing per token actually shrank even as the total expert pool more than doubled compared with Kimi K2, suggesting the efficiency claim is architectural rather than just a marketing line - a video-based observation rather than a figure drawn from Moonshot's own technical report. Community testers who ran the released weights themselves report a real-world tradeoff behind that efficiency story: even at reduced precision, the weights span roughly six GPUs in a multi-GPU server setup, so open-weight access does not mean the model runs on a single machine.


