DocsQuick StartAI News
AI NewsKimi K3 Explodes in Popularity for Three Days: Moonshot Forced to Suspend New User Subscriptions
Industry News

Kimi K3 Explodes in Popularity for Three Days: Moonshot Forced to Suspend New User Subscriptions

2026-07-19T20:05:55.220Z
Kimi K3 Explodes in Popularity for Three Days: Moonshot Forced to Suspend New User Subscriptions

Within 48 hours of Kimi K3’s release, the number of requests approached the cluster limit. Moonshot AI announced a suspension of new user subscriptions on the consumer side, prioritizing the experience of existing subscribers. Behind this “sweet dilemma” lies a new challenge in computing power scheduling for the era of open models.

On July 19, Moonshot AI (Moon’s Dark Side) issued a bittersweet announcement: effective immediately, Kimi’s consumer‑end subscriptions are suspended for new users.

It had only been 72 hours since the launch of Kimi K3.

The official reason was blunt — in the past 48 hours, user traffic had vastly exceeded expectations and approached the current cluster’s capacity limit. To preserve the experience for existing subscribers, new users had to be shut out so that all compute could serve the current base. This wasn’t throttling; it was closing the door.

Screenshot of Moonshot AI Kimi’s official announcement explaining the suspension of new subscriptions

A model that crashed its own servers

Kimi K3 went live on the evening of July 16: 2.8 trillion parameters, MoE architecture, 1 million‑token context, native visual understanding. According to Moonshot AI, this is “the world’s first open‑source model at the 3‑trillion‑parameter level,” with a larger parameter scale than the trillion‑level open‑source models previously released by DeepSeek, Zhipu GLM, and Meituan LongCat.

More important was its capability demo. The highlight of the launch materials: during a continuous 48‑hour autonomous‑agent run, K3 used open‑source EDA tools and the Nangate 45 nm process library to independently complete the design, optimization, and verification of an AI chip — 4 mm² area, 1.46 million standard cells, 0.277 MB SRAM, achieving timing closure at 100 MHz and decoding simulation throughput of over 8,700 tokens per second.

The chip itself isn’t spectacular – 45 nm is a mature process from more than a decade ago – but the significance lies in the workflow: one agent managing the entire path from RTL to timing closure. In industry, that typically takes hundreds of engineers working for years.

Add to that demos such as building a GPU programming system from scratch, generating a 3‑D open‑world game, and analyzing 391 gravitational‑wave events with 20 concurrent sub‑agents – and the dev community exploded the moment these were shown.

The ARR curve from Zhang Yutong’s Moments post

On July 18, Kimi President Zhang Yutong shared a chart in her WeChat Moments showing that on July 17, the day K3 launched, Kimi’s annual recurring revenue (ARR) suddenly spiked out of its previous flat range, marking the largest one‑day increase in history.

That curve basically explains why the system crashed.

K3’s pricing isn’t cheap — API rates:

  • Input (cache hit): ¥ 2 / million tokens
  • Input (cache miss): ¥ 20 / million tokens
  • Output: ¥ 100 / million tokens

Compared with peers: pricier than DeepSeek’s tier but roughly two‑thirds cheaper than the Claude Opus series. Positioned as “Claude Opus 4.8 performance at one‑third the price,” the value proposition is clear.

The result: consumer subscriptions flooded in far faster than expected, while developer API tests piled on traffic from the other side. The cluster maxed out completely.

Kimi K3 performance comparison chart with Claude, GPT, and other models

Splitting entitlements: a forced product overhaul

Hidden in the announcement was an important detail — Moonshot AI will henceforth sell Kimi main entitlements (Web / App / Work) and Kimi Code entitlements separately.

In plain terms: compute demand for coding scenarios is on a totally different scale from ordinary chat. Kimi Code’s long‑running agent sessions can last hours or even days, filling 1 million tokens of context as a matter of routine. Mixing coding users and chatty users in the same subscription pool lets the former drain all resources and wrecks the latter’s experience.

Cursor and Anthropic have already gone down this path – Claude Code has its own price plan, Cursor splits Pro / Business tiers. Coding agents have effectively become a distinct product category: their compute economics differ from general chat. Moonshot AI has just been pushed by reality into making that split.

Judging by the announcement’s language, Kimi Code will likely adopt a costlier, stricter‑quota model – consistent with overseas peers: developers paying for coding agents are far less price‑sensitive than consumer subscribers.

The compute dilemma of open models

The real story isn’t Kimi itself but the structural problem it exposes: the open‑weights path is reshaping the compute economics of model services.

Closed models (GPT, Claude) handle compute expansion through mature procedures — forecast traffic, pre‑purchase capacity, and allocate on demand. Backed by Azure and AWS, OpenAI and Anthropic can plan capacity months ahead.

Open‑weights releases are different. Once public, you have no idea where demand will come from:

  • Consumers: attracted by the demo
  • API users: companies and developers prototyping via the official endpoint
  • Local deployers: download weights to run themselves (not using official compute but straining Hugging Face bandwidth)
  • Third‑party hosts: cloud vendors racing to offer inference services

The first three all directly or indirectly consume Moonshot’s resources. And for a 2.8 T‑parameter model like K3, even with MoE reducing active parameters, each inference costs far more than K2. The 1 M context pushes KV cache memory to its limit.

DeepSeek also faced overload when R1 launched – but it chose to degrade rather than close: users could still enter, though responses slowed and errors spiked. Moonshot AI took the harder line – shut the gate entirely to preserve core users.

Neither strategy is inherently right or wrong. DeepSeek’s way preserved growth metrics at the cost of user experience; Moonshot’s way protected experience but missed a conversion window.

Does Kimi K3 merit the hype?

Setting the announcement aside — where does K3 really stand?

Public benchmarks place it in the third tier, behind Claude Fable 5 and GPT‑5.6 Sol, effectively the strongest open‑source model available. The Financial Times previously reported that K3 still falls short of Anthropic’s Fable (suspended for safety reasons) yet already challenges the widespread belief that Chinese models lag U.S. models by 8 to 12 months.

Technically, the most notable components are KDA (Kimi Delta Attention) and an attention‑residual mechanism. Moonshot claims K3 achieves ~2.5× better scaling efficiency than K2. If verified in the technical report, that matters more than raw parameter count – it means larger effective models at the same training cost.

Long‑horizon agent capability is K3’s strongest bet. The 48‑hour autonomous chip‑design demo signals to the market that the model is crossing from “single‑turn Q&A” to “multi‑day engineering tasks.” If that leap proves real, it rewrites AI‑application economics – imagine a month of a developer’s work compressed into a week of K3 autonomous runtime.

Of course, demos are demos; real‑world performance depends on forthcoming developer feedback. Replication efforts on SWE‑bench are already appearing on GitHub – worth watching.

Full model weights are expected by July 27, when domestic and international compute providers will likely release hosting services. For developers, this is a turning point – with the official channel overloaded, third‑party inference may be the steadier choice for now.

Architecture diagram of Kimi K3’s AI‑chip self‑design demo

A brief industry‑wide view

Zooming out, China’s AI scene in the first half of 2026 shows a clear rhythm:

  • June 30 – Meituan LongCat 2.0: 1.6 T parameters, first trillion‑parameter model trained entirely on domestic compute
  • July 2 – Zhipu GLM 5.3
  • July 6 – Tencent Hunyuan Hy3
  • July 16 – Moonshot AI Kimi K3
  • From July 17 – MiniMax M3 and H3 appeared at WAIC

Open weights have evolved from DeepSeek’s solo choice into a consensus among China’s top players. Financial Times data: Moonshot AI’s new funding round values it at US $31.5 billion; DeepSeek ~ $71 billion; Anthropic ~ $96.5 billion; OpenAI ~ $85.2 billion. The gap is narrowing – faster than anyone expected a year ago.

Marc Andreessen recently said, “Zhipu GLM‑5.2 is the first Chinese model to match flagship U.S. lab models in public tests.” That now looks conservative.

OpenAI Hub has also added Kimi K3 access, allowing seamless switching among GPT, Claude, Gemini, DeepSeek and Kimi with a single API key – and given Moonshot’s current bottlenecks, running K3 through an aggregator may actually be the more stable choice this week.

## Epilogue: a sweet headache

The pitfall Moonshot stepped into is one every open‑model player will face sooner or later. Either you degrade the experience like DeepSeek, or you close the gate like Kimi, or you stockpile compute before launch – but no one can accurately predict a SOTA model’s actual traffic curve post‑release.

The announcement’s closing line — “We sincerely apologize to friends who were full of expectation but could not get the experience they hoped for” — reads a bit helpless yet honest. At least Moonshot chose transparency: telling users “I can’t handle it,” instead of letting new signups run into endless errors.

For developers, the next few days will likely run like this:

  1. Wait for the July 27 weights release and benchmark locally or via third parties 
  2. Integrate the API through aggregators and prepare application adaptation 
  3. Watch for the technical report, especially KDA architecture details 
  4. Track when official subscriptions reopen 

Compute shortages are temporary; architectures and methodologies are long‑term. Whether Kimi K3 deserves the attention – we’ll know in a month.

## References

Related Articles

View All

Contact Us

We usually reply quickly during business hours

Scan WeChat

Support: Hub Assistant

WeChat ID: