7/20/2026, 1:02:17 PM · foundation-models

Moonshot AI Pauses Kimi K3 Subscriptions as 48-Hour Demand Surge Overwhelms GPU Capacity

Beijing-based Moonshot AI temporarily halted new subscriptions for its record-setting 2.8-trillion-parameter Kimi K3 model after unprecedented user demand pushed its Graphics Processing Unit infrastructure to its limits within two days of launch.

Launch and Immediate Overload

<cite index="2-2">Moonshot AI suspended new consumer subscriptions for its Kimi K3 model on July 19 due to overwhelming demand that strained its computing infrastructure.</cite> <cite index="10-3">The Beijing-based startup stated on X that "over the past 48 hours, demand has pushed close to the limits of our current capacity."</cite> <cite index="2-3">"To protect the experience of existing subscribers, we're temporarily pausing new subscriptions and prioritizing compute for current members," the company stated.</cite> <cite index="1-5">Moonshot added that new subscription slots would return in stages as additional computing capacity becomes available.</cite>

Model Specifications

<cite index="2-5">Kimi K3, launched on July 16, features a 2.8 trillion-parameter Mixture-of-Experts (MoE) model with a one-million-token context window and native vision capabilities.</cite> <cite index="2-6">The model is the largest open-weight system to approach the 3-trillion-parameter mark, surpassing DeepSeek V4's 1.6 trillion parameters.</cite> Under its MoE architecture, <cite index="18-1">only 16 of 896 experts are active per token under a Stable LatentMoE framework.</cite> <cite index="11-19">Together with refined training and data recipes, these structural changes yield an approximate 2.5× improvement in overall scaling efficiency compared to Kimi K2, allowing the model to convert compute into intelligence more effectively.</cite>

Kimi K3 employs aggressive quantization: <cite index="17-26,17-27">the model uses quantization-aware training (QAT) starting from the supervised fine-tuning stage — a critical distinction from post-training quantization, as the model learns to compensate for quantization error during training, resulting in significantly less quality degradation.</cite> <cite index="17-31">The MXFP4 weight format means the full 2.8-trillion-parameter model requires approximately 1.4 terabytes of weight storage — substantially less than the roughly 5.6 terabytes that FP16 weights would demand.</cite> Full model weights are scheduled for public release by July 27, 2026, per Moonshot's official technical blog.

Benchmark Performance

<cite index="11-4">While Kimi K3's overall performance still trails the most powerful proprietary models — Claude Fable 5 and GPT-5.6 Sol — it demonstrated frontier-level performance across its evaluation suite, consistently outperforming other tested models.</cite> <cite index="16-3">On the Artificial Analysis Intelligence Index, K3 scores approximately 57, placing it fourth overall, behind Claude Fable 5 (approximately 60) and GPT-5.6 Sol (approximately 59) but ahead of Claude Opus 4.8 (approximately 56).</cite> <cite index="12-8">In blind testing by AI evaluator Arena, developers preferred Kimi over every leading U.S. model for front-end coding, including Anthropic's Fable 5 and OpenAI's GPT-5.6 Sol.</cite>

<cite index="1-6,1-7">Kimi K3 is designed for coding and agent-based tasks — workloads that typically require repeated model calls and greater inference capacity, making them more expensive to operate at scale.</cite> This design orientation directly contributed to the severity of the capacity crisis.

Infrastructure Response and Structural Adjustments

In response to the demand surge, <cite index="2-8,2-9">the company announced it would split its membership system into two distinct products — "Kimi Membership" for general web, app, and office use cases, and "Kimi Code Membership" tailored for programming workflows — a change aimed at allocating computing resources more effectively.</cite>

<cite index="1-8">Although open-weight models allow users to download and customise the underlying system, analysts noted that relatively few users are likely to host a model of Kimi K3's size because of the hardware costs involved.</cite> <cite index="11-14">Moonshot itself recommends deploying Kimi K3 on supernode configurations with 64 or more accelerators.</cite>

Broader Context: Compute Constraints and Chinese AI

<cite index="1-18">U.S. export controls on advanced Nvidia chips have made access to computing power an increasingly important constraint for AI developers in China seeking to scale their models.</cite> The capacity crunch at Moonshot illustrates how the intersection of surging domestic demand and restricted access to leading semiconductor hardware creates a structural bottleneck for Chinese Large Language Model (LLM) developers, even when their models achieve competitive frontier performance.

<cite index="1-17">Alibaba, which is an investor in Moonshot, announced separately that its 2.4 trillion-parameter Qwen3.8-Max-Preview model had launched on its AI platforms ahead of a planned open-weight release</cite> — underscoring the rapid cadence of large model releases across the Chinese AI sector.

<cite index="20-16">Moonshot is reportedly raising between $1 billion and $2 billion in a new funding round that would value the company at up to $31.5 billion</cite>, and <cite index="1-11,1-12">the move comes as the company also prepares for a potential initial public offering (IPO) in Hong Kong, with two people familiar with the matter telling Reuters that Moonshot is unwinding its existing offshore corporate structure ahead of the planned listing.</cite>

Cross-references

Sources

  1. [1]
    Moonshot AI pauses Kimi K3 subscriptions as demand strains compute capacity
  2. [2]
    Moonshot Pauses Kimi K3 Signups Amid GPU Shortage - Dataconomy
  3. [3]
    Kimi K3 48-Hour Chip Design Experiment: Why Moonshot AI Autonomous EDA Pipeline Has the Entire Semiconductor Industry Nervous - Pandaily
  4. [4]
    What Is Kimi K3? China's New AI Model With 2.8 Trillion Parameters That Rivalled GPT-5.6 – Outlook Business
  5. [5]
    Markets experience new DeepSeek shock after MoonShot AI releases Kimi K3 | Fortune
  6. [6]
    Moonshot's Kimi K3 stuns AI watchers with 2.8 trillion parameters and competitive pricing
  7. [7]
    China’s Moonshot AI Unveils Kimi K3, Raising Fresh Fears in the US
  8. [8]
    Moonshot Unveils Kimi K3 AI Model, Narrowing Gap With US Rivals - Bloomberg
  9. [9]
    China's Moonshot AI Launches Kimi K3, the World's First Open 2.8T AI Model
  10. [10]
    Moonshot AI’s Kimi K3 faces GPU crunch amid US-China tech rivalry
  11. [11]
    Kimi K3 Tech Blog: Open Frontier Intelligence
  12. [12]
    China's open-weight Kimi model stuns AI world with frontier-level results
  13. [13]
    China’s Kimi K3 comes close to Fable benchmarks at one-third the token price
  14. [14]
    Moonshot AI’s New Kimi K3 Challenges U.S. Frontier Models
  15. [15]
    China Just Dropped Another Bomb on America's Frontier AI Companies
  16. [16]
    Kimi K3: Moonshot AI’s 2.8T Open-Weight Model — Release, Specs & Pricing (2026)
  17. [17]
    Kimi K3 Model Overview: 2.8T Parameters, MXFP4 Quantization, and What the Open Weights Mean for the Community
  18. [18]
    Kimi K3 Release: Open Frontier Intelligence at 2.8T Scale
  19. [19]
    Kimi K3: The Open Model Closing the Gap | BenchLM.ai
  20. [20]
    Kimi K3: Moonshot's 2.8T Open-Weight Model Explained