ByteDance, the company behind TikTok, is training an artificial intelligence model with as many as 10 trillion parameters, according to a Financial Times report that cited three people familiar with the plan. The scale would put it among the largest models any company has attempted, and it is the clearest sign yet that China's biggest internet firms are ready to spend heavily to close the gap with the leading Western labs. The news was picked up more widely by The Next Web and Slashdot.

For context, 10 trillion parameters is more than three times the size of Moonshot's Kimi K3, which at about 2.8 trillion sits near the top of the current crop of Chinese models. The FT reported that ByteDance is aiming the new system at Anthropic's Mythos, one of the frontier models Chinese developers have so far struggled to match. Some accounts put the training effort at around 30,000 GPUs.

Why the headline number tells you less than it seems

A parameter count is an easy thing to print and a hard thing to interpret. The reporting does not say whether ByteDance is building a dense model, where every parameter is used for every word, or a mixture-of-experts model, which can hold a huge number of total parameters while only switching on a small slice of them for any given token. That distinction changes what the number means. A 10 trillion parameter mixture-of-experts system might activate a fraction of that at run time, which is why raw size is a poor guide to how capable a model actually is.

The pieces that would let anyone judge the model are still missing: the active-parameter count, the training budget, the data it learns from, and how it scores on real tests. Until those arrive, the 10 trillion figure is a statement of ambition more than a measure of quality.

Part of a wider push

The project fits a run of releases from Chinese labs that keep narrowing the distance to the frontier. In recent weeks an open Chinese model reached near-frontier scores, and DeepSeek shipped a cheaper, faster version of its flagship aimed at agent tasks. ByteDance is taking the opposite tack from that efficiency drive, betting that sheer scale still buys an edge. Pretraining a model this large usually runs three to six months, so the first real evidence, or the first signs of trouble, may not surface until late in the year.

Sources

  1. i. thenextweb.com
  2. ii. slashdot.org
  3. iii. aiweekly.co

Commentarii · 0

Add · a · Comment