The problem with releasing a model everyone wants is that everyone shows up at once. Moonshot AI found that out this week. On 20 July the Chinese lab paused new subscriptions to Kimi K3, its 2.8-trillion-parameter open model, after demand in the space of 48 hours pushed its computing clusters close to their limit.
The company was candid about it. 'Kimi K3 has received far more love than we expected, and our GPUs are feeling it,' it wrote on its official account, adding that it was pausing new signups 'to protect the experience of existing subscribers' while it added capacity 'as fast as we can' and reopened spots in batches. Anyone already subscribed keeps working as normal.
The crush follows a fortnight in which K3 has been hard to ignore. We covered its arrival as the largest open-weight model yet, and then its jump to first place on a major coding leaderboard, ahead of closed models from far better-funded rivals. That combination, strong benchmarks and open weights, is exactly the sort of thing that sends developers rushing to try something the same afternoon they read about it.
A capacity problem, not a demand problem
Alongside the pause, Moonshot said it is splitting its subscription into two tiers: a Kimi Membership covering web, app and work features, and a separate Kimi Code Membership for programming. The move reads as an attempt to steer the heaviest coding workloads, the ones eating the most compute, into their own lane.
There is an irony here that the open-weights community will appreciate. Moonshot has promised to release K3's weights on 27 July, which means the very users being turned away from the hosted service will soon be able to run the model themselves, capacity permitting. For now the bottleneck is the same one throttling everyone in this business, from Shanghai to San Francisco: not ideas, but GPUs.
Sources
- i. the-decoder.com
- ii. www.pymnts.com
- iii. www.euronews.com
Commentarii · 0