Launch and Immediate Overload
<cite index="2-2">Moonshot AI suspended new consumer subscriptions for its Kimi K3 model on July 19 due to overwhelming demand that strained its computing infrastructure.</cite> <cite index="10-3">The Beijing-based startup stated on X that "over the past 48 hours, demand has pushed close to the limits of our current capacity."</cite> <cite index="2-3">"To protect the experience of existing subscribers, we're temporarily pausing new subscriptions and prioritizing compute for current members," the company stated.</cite> <cite index="1-5">Moonshot added that new subscription slots would return in stages as additional computing capacity becomes available.</cite>
Model Specifications
<cite index="2-5">Kimi K3, launched on July 16, features a 2.8 trillion-parameter Mixture-of-Experts (MoE) model with a one-million-token context window and native vision capabilities.</cite> <cite index="2-6">The model is the largest open-weight system to approach the 3-trillion-parameter mark, surpassing DeepSeek V4's 1.6 trillion parameters.</cite> Under its MoE architecture, <cite index="18-1">only 16 of 896 experts are active per token under a Stable LatentMoE framework.</cite> <cite index="11-19">Together with refined training and data recipes, these structural changes yield an approximate 2.5× improvement in overall scaling efficiency compared to Kimi K2, allowing the model to convert compute into intelligence more effectively.</cite>
Kimi K3 employs aggressive quantization: <cite index="17-26,17-27">the model uses quantization-aware training (QAT) starting from the supervised fine-tuning stage — a critical distinction from post-training quantization, as the model learns to compensate for quantization error during training, resulting in significantly less quality degradation.</cite> <cite index="17-31">The MXFP4 weight format means the full 2.8-trillion-parameter model requires approximately 1.4 terabytes of weight storage — substantially less than the roughly 5.6 terabytes that FP16 weights would demand.</cite> Full model weights are scheduled for public release by July 27, 2026, per Moonshot's official technical blog.
Benchmark Performance
<cite index="11-4">While Kimi K3's overall performance still trails the most powerful proprietary models — Claude Fable 5 and GPT-5.6 Sol — it demonstrated frontier-level performance across its evaluation suite, consistently outperforming other tested models.</cite> <cite index="16-3">On the Artificial Analysis Intelligence Index, K3 scores approximately 57, placing it fourth overall, behind Claude Fable 5 (approximately 60) and GPT-5.6 Sol (approximately 59) but ahead of Claude Opus 4.8 (approximately 56).</cite> <cite index="12-8">In blind testing by AI evaluator Arena, developers preferred Kimi over every leading U.S. model for front-end coding, including Anthropic's Fable 5 and OpenAI's GPT-5.6 Sol.</cite>
<cite index="1-6,1-7">Kimi K3 is designed for coding and agent-based tasks — workloads that typically require repeated model calls and greater inference capacity, making them more expensive to operate at scale.</cite> This design orientation directly contributed to the severity of the capacity crisis.
Infrastructure Response and Structural Adjustments
In response to the demand surge, <cite index="2-8,2-9">the company announced it would split its membership system into two distinct products — "Kimi Membership" for general web, app, and office use cases, and "Kimi Code Membership" tailored for programming workflows — a change aimed at allocating computing resources more effectively.</cite>
<cite index="1-8">Although open-weight models allow users to download and customise the underlying system, analysts noted that relatively few users are likely to host a model of Kimi K3's size because of the hardware costs involved.</cite> <cite index="11-14">Moonshot itself recommends deploying Kimi K3 on supernode configurations with 64 or more accelerators.</cite>
Broader Context: Compute Constraints and Chinese AI
<cite index="1-18">U.S. export controls on advanced Nvidia chips have made access to computing power an increasingly important constraint for AI developers in China seeking to scale their models.</cite> The capacity crunch at Moonshot illustrates how the intersection of surging domestic demand and restricted access to leading semiconductor hardware creates a structural bottleneck for Chinese Large Language Model (LLM) developers, even when their models achieve competitive frontier performance.
<cite index="1-17">Alibaba, which is an investor in Moonshot, announced separately that its 2.4 trillion-parameter Qwen3.8-Max-Preview model had launched on its AI platforms ahead of a planned open-weight release</cite> — underscoring the rapid cadence of large model releases across the Chinese AI sector.
<cite index="20-16">Moonshot is reportedly raising between $1 billion and $2 billion in a new funding round that would value the company at up to $31.5 billion</cite>, and <cite index="1-11,1-12">the move comes as the company also prepares for a potential initial public offering (IPO) in Hong Kong, with two people familiar with the matter telling Reuters that Moonshot is unwinding its existing offshore corporate structure ahead of the planned listing.</cite>