Alibaba released Qwen 3.8 Max on August 3, and the company is not being shy about where it thinks the model belongs. On its own published benchmarks, the new system edges past OpenAI's GPT-5.6 Sol and Anthropic's Claude Fable 5 on several tests, positioning a Chinese lab squarely against the two names that have defined the frontier this year.

The model is a 2.4-trillion-parameter mixture-of-experts design, the largest Alibaba has shipped. It accepts text, images and video, returns text, and carries a context window of up to one million tokens. Pricing lands at $2.00 per million input tokens and $6.00 per million output, with cached input far cheaper at $0.25, which undercuts the flagship tiers from the American labs by a wide margin.

The benchmark case

Alibaba's strongest claim is on PaperBench, which measures whether a model can reproduce the results of a research paper. Qwen 3.8 Max scores 93.0 there, ahead of GPT-5.6 Sol at 90.5, Fable 5 at 88.8, and Claude Opus 4.8 at 80.3. On Terminal-Bench 2.1, a test of agent behaviour in a command line, it posts 86.6. That sits above Opus 4.8 and Fable 5, both at 84.6, though still behind GPT-5.6 Sol's 88.8.

The company also points to sharp gains in multimodal reasoning over the previous Qwen-Max generation. Those figures are worth reading with care. Every number so far comes from Alibaba itself. Independent evaluators had not published their own scores at release, and vendor benchmarks have a long history of flattering the model that ships them. The real test will be how Qwen 3.8 Max holds up on leaderboards it does not control.

The open-weights promise

What sets this launch apart is the plan to release the weights. Alibaba says it will publish an open version within a week, along with a smaller companion checkpoint, Qwen 3.8 27B, aimed at teams that want to run a model on their own hardware rather than call an API. A Max-class Qwen model has never been released openly before. If it happens on schedule, it would put a system with frontier-level claims into anyone's hands for free.

That prospect lands at a pointed moment. Washington has been weighing curbs on Chinese open-weight models precisely because their quality and low cost have made them hard to ignore. Qwen 3.8 Max is the clearest example yet of why that debate exists. It arrives days after Google shipped its own Gemini Flash trio, and the pace shows no sign of easing.

For now the honest verdict is a split one. The pricing is real and aggressive. The openness, if it holds, matters. The benchmark supremacy is a claim, not yet a finding, and it deserves the scrutiny that any lab grading its own homework should expect.

Sources

  1. i. www.marktechpost.com
  2. ii. www.neowin.net
  3. iii. officechai.com

Commentarii · 0

Add · a · Comment