Every few weeks a new model arrives waving a bigger number. This month it was Alibaba's 2.4-trillion-parameter Qwen3.8 Max and Moonshot's 2.8-trillion-parameter Kimi K3. The implied promise is simple and seductive. More parameters, more intelligence. It is also, as a rule of thumb, wrong.
Where the idea came from
The belief had a good run. For several years, making models larger really did make them better, and the scaling laws that described this became something close to industry gospel. Parameter count turned into a marketing figure, the horsepower rating of AI. The trouble is that the relationship was never as clean as the slogans suggested, and it began breaking down in public some time ago.
The clearest early crack came from DeepMind's own work. Its 70-billion-parameter Chinchilla model outperformed the 280-billion-parameter Gopher, on the same compute budget but with far more training data. The lesson was not that size hurts. It was that size without enough good data to match it is wasted. A smaller model, better fed, won.
What actually moves the needle
Since then the field has found several routes for smaller models to beat larger ones. Better training data is one. Smarter architecture is another. A newer one is letting a model spend more effort at the moment you ask a question, so a compact model that "thinks" for longer can outscore a giant that answers on reflex. For plenty of ordinary tasks the tiny models already top out the benchmark. A sub-billion-parameter model can hit well over ninety percent on simple text classification, and the enormous version barely improves on it.
There is also the awkward fact that a headline parameter count tells you almost nothing you can act on. Most of the largest models are now mixture-of-experts designs, where only a slice of those parameters is active on any given request. A "2.4 trillion parameter" banner and the number actually doing the work are two very different figures, and the banner is the one that reaches the press release. Headline scores are notoriously easy to game, and parameter counts even more so.
How to read the next big number
None of this means the new Chinese giants are weak. Some may turn out excellent. The point is narrower and more useful. A parameter count is not a quality score, and you should not treat it as one. What matters is how a model performs on the work you actually care about, ideally measured by someone who did not build it. So when the next multi-trillion-parameter flagship lands with a record number attached, the honest response is not to be impressed. It is to ask what it can really do, and to wait for someone independent to check.
Sources
- i. medium.com
- ii. fonzi.ai
- iii. arxiv.org
- iv. www.seangoedecke.com
Commentarii · 0