When we covered the launch of Kimi K3 earlier this week, the interesting number was its size. Moonshot AI had shipped a 2.8-trillion-parameter model and promised to give the weights away. What it had not yet done was prove the thing could compete at the top of the field.
The benchmark results have now arrived, and they are better than most people expected.
First place, and not by a rounding error
K3 took the number one position on Arena.AI's Frontend Code Arena with a score of 1,679. Claude Fable 5 scored 1,631. GPT-5.6 Sol scored 1,618. In a field where leaderboard positions usually change hands by a point or two, a 48-point gap over the nearest closed model is a real margin.
The pattern holds across other coding evaluations. On SWE Marathon, a benchmark built around long-running software engineering tasks, K3 scored 42.0 against Fable 5's 35.0 and GPT-5.6 Sol's 39.0. On Program Bench it scored 77.8, narrowly ahead of Sol at 77.6 and Fable 5 at 76.8.
K3 does not win everything. On Moonshot's own internal Kimi Code Bench 2.0 it scores 72.9, behind Fable 5 at 76.9. A model scoring second on its maker's in-house test while leading the independent arenas is an unusual result, and an honest one to publish.
What the architecture is doing
The headline parameter count is misleading on its own. K3 uses a mixture-of-experts design that activates just 16 of its 896 experts for any given token, roughly 1.8 percent of the pool. The model is enormous in storage and comparatively cheap to run.
That sparsity is the reason the pricing works. Moonshot lists the API at $3 per million input tokens and $15 per million output tokens, which sits under most Western frontier pricing rather than dramatically below it. The context window is 1 million tokens, about four times what K2.6 handled, and the model takes text, images and video.
The date that matters is 27 July
Right now K3 is available through Moonshot's own products and its API. There is still no public model card and no weights download. Moonshot has said the full weights arrive on 27 July under a modified MIT licence, which would make this the most capable model anyone can self-host.
That is the part worth watching. Benchmark scores move around and leaderboards get gamed. A set of downloadable weights that runs on your own hardware is a different kind of fact, and it does not expire when the next model ships.
The timing is not accidental either. Anthropic spent the past fortnight reshuffling who gets Fable 5 access and at what price, and Google's flagship is still not out. A Chinese lab publishing competitive weights into that gap is a reasonable read of the moment.
Whether the weights actually appear on 27 July is the open question. Moonshot has already let one soft deadline pass without a model card. We will cover it when the files are up.
Sources
- i. venturebeat.com
- ii. www.tomshardware.com
- iii. felloai.com
- iv. openrouter.ai
Commentarii · 0