The open-weight field keeps closing the gap with the labs that keep their models locked up. On 1 June, the Chinese company MiniMax released M3, a model it describes as the first to combine frontier-level coding, a one-million-token context window and native image and video understanding in a single open package.

The interesting part is underneath. M3 is built on a new design the company calls MiniMax Sparse Attention. According to a technical write-up from MarkTechPost, the architecture cuts the per-token compute at a million tokens of context to a twentieth of the prior generation, with prefill more than nine times faster and decoding more than fifteen times faster. Long context has been expensive to run at scale, and cost is exactly where open models tend to win converts.

The benchmark claims, and the caveat

On coding, MiniMax reports a score of 59.0 percent on SWE-Bench Pro, which it says beats GPT-5.5 and Gemini 3.1 Pro and comes within reach of Claude Opus 4.7. VentureBeat frames the appeal as frontier-class results at five to ten percent of the cost.

Those numbers deserve a note of caution. As Tech Times pointed out at launch, the figures are the company's own and have not yet been independently verified. The open weights and the full technical report are due on Hugging Face and GitHub within about ten days of release, which is when outside researchers can start checking the claims for themselves.

M3 lands in a busy stretch for open releases. In the past week alone, NVIDIA has shipped Nemotron 3 Ultra and the Cosmos 3 foundation model for physical AI. The pattern is hard to miss. The gap between what you can download and what you have to rent is narrowing month by month, and a good deal of the pressure is now coming from outside the United States.

Sources

  1. i. www.minimax.io
  2. ii. www.marktechpost.com
  3. iii. venturebeat.com
  4. iv. www.techtimes.com
  5. v. pandaily.com

Commentarii · 0

Add · a · Comment