Moonshot AI released Kimi K3, a 2.8-trillion-parameter open-weight Mixture-of-Experts model with a 1M-token context window, the same week it closed a $3.5 billion funding round at a $35 billion valuation ahead of a planned Hong Kong IPO.
TECH

Moonshot AI released Kimi K3, a 2.8-trillion-parameter open-weight Mixture-of-Experts model with a 1M-token context window, the same week it closed a $3.5 billion funding round at a $35 billion valuation ahead of a planned Hong Kong IPO.

36+
Signals

Strategic Overview

  • 01.
    Moonshot AI released Kimi K3, a 2.8-trillion-parameter open-weight Mixture-of-Experts model with 104 billion activated parameters, native visual understanding, and a 1,048,576-token (1M) context window.
  • 02.
    The model runs on a hybrid architecture - 69 Kimi Delta Attention (KDA) linear-attention layers plus 24 Gated MLA layers, with 896 experts and 16 active per token.
  • 03.
    Model weights went public on July 27, 2026, letting developers download, modify, and self-host the model for free, after Moonshot previewed the model's benchmark claims around July 16.
  • 04.
    Demand surged so fast after launch that Moonshot's GPU capacity hit its limits within 48 hours, forcing a temporary pause on new subscriptions to protect existing users.
  • 05.
    The launch landed the same week Moonshot closed a $3.5 billion funding round at a $35 billion valuation - beating its original $1-2 billion target - led by China's National Artificial Intelligence Industry Investment Fund, ahead of a planned Hong Kong IPO.

Inside the Architecture: How KDA and MLA Deliver 2.5x Efficiency

Kimi K3's headline number isn't just its 2.8 trillion total parameters - it's how few of them fire on any given token. The model activates 104 billion parameters through a mixture of 896 experts with only 16 active per token, spread across 93 layers [1]. Rather than stacking more standard attention layers, Moonshot built a hybrid stack of 69 Kimi Delta Attention (KDA) linear-attention layers and 24 Gated MLA layers, a design SGLang and the Miles RL training framework backed with simultaneous day-zero support [2]. On SGLang's disaggregated serving setup, the model reached 423 tokens per second at batch-1 decode with speculative decoding and 2,808 tokens per second per GPU [2]. The efficiency case matters because it's also winning benchmarks outright - Kimi K3 took the #1 spot on the Frontend Code Arena at 1,679 Elo, ahead of Claude Fable 5 and GPT-5.6 Sol, and posted the best open-weight GPQA Diamond score ever published [1][3]. It still trails the very top closed models on Artificial Analysis's broader Intelligence Index, landing 4th overall - a reminder that 'best open-weight' and 'best overall' aren't yet the same claim [3].

The Money Story: A $35 Billion Round and a Hong Kong IPO on Deck

Kimi K3's release wasn't a standalone product moment - it landed the same week Moonshot closed a $3.5 billion funding round at a $35 billion valuation, blowing past its original $1-2 billion target [4]. That number caps an extraordinary run: Alibaba led Moonshot's first major round at a $2.5 billion valuation in February 2024, a follow-on pushed it to roughly $3 billion by the end of that year, and it sat near $4.3 billion at the end of 2025 before rocketing to $20 billion in May 2026 and now $35 billion [5][6]. Moonshot is reportedly already back in front of investors seeking a $50 billion pre-money valuation ahead of a planned Hong Kong IPO, with Goldman Sachs and CICC said to be in talks for lead underwriter roles [7][8]. The fundamentals back some of that momentum - annual recurring revenue reportedly rose from $200 million in April 2026 to $300 million by June [7]. What's notable about the investor list is China's National Artificial Intelligence Industry Investment Fund leading the latest round - the same state vehicle that also backs DeepSeek, turning Moonshot's fundraising into a visible marker of Beijing's direct financial support for open-weight AI as a strategic export.

Washington's Open-Weight Reckoning

Kimi K3 didn't just move technical benchmarks - it moved markets and reopened a policy fight. The Nasdaq fell roughly 1% after the release as investors sold Intel and Nvidia shares, a reaction analysts tied to anxiety that Kimi K3 had narrowed the US-China frontier AI gap to something like three to six months [10][11]. That anxiety split Silicon Valley and Washington along unusual lines. OpenAI's Dean Ball, a former senior AI advisor to President Trump, expressed surprise that the Chinese state keeps allowing models this capable to be open-sourced and warned of coming regulatory scrutiny, calling the dynamic 'AI communism' [9]. Venture capitalists pushed back hard: David Sacks accused the leading closed labs of lobbying to use regulation to kill open-source competition, while Chamath Palihapitiya argued the industry should simply embrace open source rather than resist it [9]. Others reject the premise that this is imitation rather than innovation - Stanford's Graham Webster pointed to 'real innovation going on' inside Chinese labs [10]- while Nvidia's Jensen Huang said US companies 'absolutely should be allowed to use Chinese models' [13]. The result is an unresolved clash over whether to restrict Chinese open-weight models by policy [12].

Open Weights, Closed Doors: Who Can Actually Run This

The 'open' in open-weight is doing a lot of work here. Commercial infrastructure moved fast - Baseten shipped a day-zero Model API with vision input and the full 1M-token context on NVIDIA GB300 NVL72 systems, building a custom tokenizer up to 18x faster for long sequences, while vLLM published its own deployment recipes co-developed with Moonshot, NVIDIA, and AMD [14][15]. That kind of hosted-API convenience is exactly why the open-weight framing is being scrutinized: analysis of the launch points out that open weights lower adoption friction enough that production users like DoorDash, Coinbase, and Cursor have integrated Kimi models, and Chinese models now reportedly account for roughly 60% of US token usage on OpenRouter [13]. But downloading the weights and actually self-hosting them are two different propositions - developers who tried it found the full model requires roughly 1.4 terabytes of storage and enterprise-grade GPU clusters, well beyond what almost any individual or small team can field, which pushes most usage back toward hosted Chinese API providers rather than genuine independent control over inference and data. Enthusiasts have gotten creative regardless, streaming quantized experts on demand to run a version of the model on consumer hardware - proof the weights are real, even if 'open' doesn't yet mean 'accessible.'

Historical Context

2024-02-01
Alibaba led Moonshot's first major funding round, valuing the company at $2.5 billion post-money on a $1 billion raise.
2024-12-01
A follow-on round pushed Moonshot's valuation to roughly $3 billion.
2025-11-06
Alibaba-backed Moonshot released Kimi K2 Thinking, an earlier model in the Kimi line.
2026-02-17
Moonshot sought a $10 billion valuation in a new funding round.
2026-05-07
Moonshot raised funds at roughly a $20 billion valuation, a near-fivefold jump from about $4.3 billion at end of 2025.
2026-07-16
Moonshot unveiled Kimi K3 with benchmark claims rivaling Anthropic and OpenAI's top models, ahead of the July 27 open-weight release.
2026-07-27
Kimi K3 model weights were released for public download, becoming what BBC described as the world's first open-source model in the three-trillion-parameter class.
2026-07-29
Moonshot closed a $3.5 billion funding round at a $35 billion valuation, exceeding its original $1-2 billion target, with China's National AI Industry Investment Fund among lead investors.

Power Map

Key Players
Subject

Moonshot AI released Kimi K3, a 2.8-trillion-parameter open-weight Mixture-of-Experts model with a 1M-token context window, the same week it closed a $3.5 billion funding round at a $35 billion valuation ahead of a planned Hong Kong IPO.

BA

Baseten

Inference infrastructure provider that shipped day-zero Model APIs for Kimi K3 with vision input and full 1M-token context, running on NVIDIA GB300 NVL72 systems; collaborated with vLLM and SGLang teams on benchmarking and built a custom tokenizer up to 18x faster for long sequences.

SG

SGLang (LMSYS)

Open-source inference engine that announced simultaneous day-0 support with the 'Miles' RL training framework, co-developed with Moonshot AI and NVIDIA; achieved 423 tokens/second batch-1 decode and 2,808 tokens/second per GPU in disaggregated serving.

VL

vLLM

Shipped day-zero support with published deployment recipes, developed alongside Moonshot AI, NVIDIA, and AMD.

AL

Alibaba Cloud and Huawei (Ascend)

Provided day-zero Kimi K3 support on their respective cloud/accelerator ecosystems; Alibaba is also Moonshot's largest outside investor, holding roughly a 36% stake dating to a February 2024 lead round.

NA

National Artificial Intelligence Industry Investment Fund (China state fund)

Lead investor in Moonshot's $3.5 billion round; the same state vehicle also backs DeepSeek, illustrating Beijing's direct financial support for the open-weight AI push.

GO

Goldman Sachs and CICC (China International Capital Corporation)

Reportedly in discussions for lead underwriter roles on Moonshot's planned Hong Kong IPO.

DO

DoorDash, Coinbase, Cursor

Cited as production users that have integrated Kimi models into their workflows, illustrating commercial adoption of the open-weight model.

Fact Check

15 cited
  1. [1] Kimi-K3 Model Card
  2. [2] Kimi K3 Day-0 Support
  3. [3] Kimi K3 Benchmarks Comparison 2026
  4. [4] Moonshot AI Reaches $35B Valuation After $3.5B Funding Round
  5. [5] Moonshot Kimi: A History
  6. [6] Chinese AI Model Developer Kimi Raising Funds Valuing It at $20 Billion
  7. [7] Moonshot AI Eyes Hong Kong IPO as China AI Race Heats Up
  8. [8] Moonshot AI Eyes $50 Billion Valuation
  9. [9] Silicon Valley Is Freaking Out Over China's Open-Source AI Strategy
  10. [10] Kimi K3: The Moonshot Chinese AI Firm Rivaling Anthropic and OpenAI
  11. [11] Why China's Open-Weight AI Model Kimi K3 Is Sparking Anxiety in Silicon Valley
  12. [12] China's Open-Source AI Model Sparks Policy Clash in Washington
  13. [13] Why Kimi K3 Signals a Convergence Toward Open-Weight Models
  14. [14] How to Build a Day-Zero API for Kimi K3
  15. [15] vLLM Day-Zero Support for Kimi K3

Source Articles

Top 5

THE SIGNAL.

Analysts

Expressed surprise the Chinese state continues allowing open-sourcing of models this good and warned of coming regulatory risk around use of open-weight Chinese models, framing open models as potentially destabilizing. Quote: "AI communism"

Dean Ball
OpenAI Head of Strategy; former senior AI advisor to President Trump

Criticized closed AI labs for lobbying to eliminate open-source competition through regulation. Quote: "The leading closed labs, already a duopoly in terms of AI model revenue, want the government to eliminate their open source competition"

David Sacks
Venture capitalist, co-chair of the President's Council of Advisors on Science and Technology

Argued the industry should embrace open-source AI rather than resist it. Quote: "The future is open source. We need to embrace it and get on with it"

Chamath Palihapitiya
Venture capitalist, All-In Podcast co-host

Argued China's AI advances reflect genuine technical innovation, not just distillation of US models. Quote: "real innovation going on"

Graham Webster
Professor, Stanford University

Voiced support for competitive access to Chinese AI models for US companies. Quote: "absolutely should be allowed to use Chinese models"

Jensen Huang
CEO, Nvidia
The Crowd

We compared 1-bit Kimi K3 to Claude Opus 5 and GPT 5.6. We gave 4 models the same prompt: Create a glass aquarium whose side panel develops a visible crack and then bursts. 1-bit Kimi K3 GGUF ran locally on 4x B200s at 36 tok/s. GitHub repo: https://t.co/aZWYAtakBP https://t.co/sGjouOoDkU

@@UnslothAI445

An engineer just spent 48 hours on reverse-engineering Kimi K3's entire codebase. What he found is wild. K3 has 2.8 trillion parameters. That's 22,580 GPT-2s packed into one model. 22,000x growth in seven years. But the real insight isn't the size. It's that every

@@VaibhavSisinty272

One open-source model just took #1 on the Frontend Code Arena and beat Claude Fable, GPT-5.6, and every closed flagship. Kimi K3 from Moonshot AI. 2.8 trillion parameters. Previous open-source king DeepSeek sat at 1.6T. 1 million token context. Reads full codebases and

@@Mayaikos20

Got Kimi K3 running on my MacBook. It's painfully slow, but it works.

@u/gavanon383
Broadcast
Kimi K3 explained in 13min..

Kimi K3 explained in 13min..

Build Anything with Kimi K3, Here’s How

Build Anything with Kimi K3, Here’s How

Kimi K3 Is Fable Level... (they should be worried)

Kimi K3 Is Fable Level... (they should be worried)