A Frontier Benchmark Win That Broke Its Own Servers
Kimi K3 went live on July 16, 2026 across the Kimi app, Kimi Work, Kimi Code, and the Kimi API, timed to land just ahead of the World Artificial Intelligence Conference in Shanghai, with full open weights scheduled to follow on July 27 [3]. The model is a 2.8-trillion-parameter Mixture-of-Experts system - 896 experts, 16 active per token - built on a hybrid linear-attention design called Kimi Delta Attention, with native vision support and up to a 1 million token context window [2]. It immediately took the top spot on LMArena's Frontend Code Arena leaderboard with a score of 1,679, edging out Claude Fable 5 and GPT-5.6 Sol [4], posting an overall Elo of 1547 - a jump of 732 points over predecessor Kimi K2.6 [1].
The win came with an immediate operational cost. Within 48 hours of launch, user demand pushed Moonshot's GPU capacity to its limit, forcing the company to pause new subscriptions while protecting existing users and preparing to split membership into separate general-use and coding tiers [15]. That gap between benchmark supremacy and infrastructure readiness is the quieter story here: a lab that just out-coded the best US models could not simultaneously serve the audience its win attracted, a reminder that frontier capability and frontier operational capacity are not the same achievement.



