On July 19, Moonshot AI temporarily suspended new subscriptions for its generative AI service "Kimi." According to an official X post, demand for the latest model, "Kimi K3," approached the company's current capacity ceiling within the past 48 hours. Existing subscribers are unaffected, and the company says it will prioritize allocating compute resources to current members.

What's at issue here isn't whether the model can be provided at all. Immediately after demand for K3 surged, the question of how to distribute inference GPUs across different entry points—Web, app, work support, and coding—has entered a phase where it directly determines service quality. Moonshot has also indicated plans to add capacity and gradually reopen new subscription slots.

AD

Halting New Subscriptions to Redirect Compute to Existing Members

The suspension applies to new subscriptions, and Moonshot states that existing subscribers will not be affected. This is a decision to prioritize users who are already paying, at a moment when demand is approaching capacity limits. The company also noted it would reopen new slots gradually in line with added capacity, rather than restoring them all at once—a sign that it intends to manage the situation based on available surplus.

However, the announcement did not specify a resumption date, the current number of subscribers, or the scale of GPU additions. Pricing and usage caps were also not disclosed, leaving unclear how much K3 usage each member will be able to secure. The announcement pertains specifically to new membership subscriptions and does not address Kimi API availability or rate limits. The suspension of new subscriptions should not be read as equivalent to API restrictions.

Running a 2.8-Trillion-Parameter-Class Model Across Multiple Pathways

K3 has a total of 2.8 trillion parameters, native visual understanding, and a 1-million-token context window. Moonshot states it is available through Kimi.com, Kimi Work, Kimi Code, and the Kimi API—meaning demand for the same model flows in through a wide range of entry points, from conversational use to external applications. The current membership suspension is a measure that narrows just one of these entry points: new membership subscriptions.

The model uses a Mixture of Experts (MoE) architecture, activating 16 of 896 expert modules at runtime. While not all parameters are engaged on every pass, Moonshot recommends a supernode configuration bundling 64 or more accelerators to run inference efficiently. This does not necessarily reflect the actual configuration the company is using. Still, for models like K3 that are built around long context and agentic work, the design of infrastructure for stable inference delivery remains a challenge separate from the model's own release policy.

AD

Splitting Membership Plans by Use Case

Going forward, Moonshot will offer a Kimi Membership centered on Web usage, which will also cover the app and Kimi Work. Coding workflows will be handled by a separate Kimi Code Membership. The official post states that this two-tier split is intended to allocate compute resources more precisely and stabilize the user experience. In the same post announcing the subscription suspension, the company also signaled a shift toward allocating capacity differently depending on use case.

However, pricing, usage limits, migration terms for existing members, and a rollout timeline for the new plans have not yet been announced. Kimi Code allows users to select K3 for handling long codebases, and K3 is also available through Web and Work. Under current Kimi Code documentation, K3 is available to members at the Moderato tier or above, while the 1-million-token context is reserved for Allegretto tier or above. It has not been announced how the two new membership tiers will carry over these existing tiers and usage limits.

GPU Constraints Remain in Hosted Inference

The 2.8-trillion-parameter-class K3 is offered through Kimi.com, Kimi Work, Kimi Code, and the Kimi API. Because the responses users receive from Kimi are generated in a hosted environment, an increase in available access pathways does not automatically mean an increase in the service's inference capacity. This suspension has, in a short span of time, made clear that the scope of a model's availability and the capacity needed to reliably deliver responses are separate operational issues.

If the capacity expansion and phased reopening proceed as planned, the subscription halt will amount to an operational measure absorbing the demand spike immediately following K3's release. On the other hand, if the usage limits of the new plans or the frequency of reopenings turn out to reflect an ongoing constraint, then choosing K3 will require looking not just at performance, but at which pathway can secure which compute resources. What Moonshot should disclose next isn't the scale of the added GPUs, but subscription terms and resumption conditions that users can actually predict.