Moonshot Kimi K3 pauses subscriptions 48h post-launch as demand surges sixfold; 2.8T-param model strains GPU allocation
Chinese AI startup Moonshot temporarily halted new subscription onboarding for its Kimi K3 model on July 19, less than 48 hours after launch, as user demand overwhelmed GPU compute capacity. The company announced the pause on X on Sunday (July 19), stating demand had 'pushed close to the limits of our current capacity.' Kimi K3, launched around July 16-17, is a 2.8-trillion-parameter mixture-of-experts open-weight model with a 1-million-token context window, designed for long-horizon coding, knowledge work, and agentic reasoning tasks.
Third-party benchmarks fueled adoption frenzy: Arena ranked Kimi K3 first for web interface building, outperforming Claude Fable 5 and OpenAI's GPT-5.6 Sol on specific developer workflows. Moonshot reported a sixfold surge in demand within days of launch. The company stated it would reopen subscription slots in batches as infrastructure capacity expands, protecting existing paid subscribers from degradation. Moonshot also restructured membership plans into Kimi Membership (web, app, work) and Kimi Code Membership (coding workflows).
The demand spike arrives as Moonshot pursues a Hong Kong IPO—advisors include Goldman Sachs and CICC—and seeks $2 billion in fresh capital. The company's valuation has climbed to $30 billion as of June 2026. The startup reported $300 million ARR driven by API demand, signaling underlying monetization is working despite the cloud-capacity crunch. Full weights for Kimi K3 are scheduled for public release on July 27 under a Modified-MIT license, at which point distributed self-hosting becomes an option for those with sufficient hardware.
For architects in China: Kimi K3's subscription pause crystallizes the structural constraint facing Chinese AI development—US export controls restrict NVIDIA H100 and Blackwell access, forcing reliance on older chips and expensive parallel capacity. Moonshot recommends serving K3 on supernodes of at least 64 accelerators; the model requires 1.5TB GPU memory full precision or 600GB INT4 quantized. The open-weight release (July 27) shifts hosting burden to global developer community, but infrastructure economics remain tight. This demand surge validates Chinese open-weight strategy but underscores capex intensity as the binding constraint in the next phase of competition.
Sources
- Primary source
- pymnts.com
“Kimi K3 received far more love than expected; demand pushed GPUs close to limits within 48 hours”
- easternherald.com
“Demand surged sixfold within days; Moonshot pursuing $30B Hong Kong IPO”
- finance.yahoo.com
“Kimi K3 is 2.8T-parameter open-weight model; ARR $300M driven by API demand; full weights release July 27”