A Chinese lab just took the top spot on the most-watched AI coding benchmark, using hardware the US tried to keep out of China's hands. Z.ai, formerly Zhipu AI, released GLM-5.1 on April 7, 2026 under an MIT license. It scored 58.4 on SWE-Bench Pro, edging past GPT-5.4 at 57.7 and Claude Opus 4.6 at 57.3. It's the first time an open-source model has topped a major real-world software engineering leaderboard.
What the model is
GLM-5.1 uses a Mixture-of-Experts architecture with 754 billion total parameters, though only 40 billion are active for any given input. That keeps inference costs manageable despite the model's overall size. It supports a 200,000-token context window and can generate up to 128,000 tokens in a single response. The full weights are on HuggingFace with no commercial restrictions.
Z.ai spun out of Tsinghua University and has been developing the GLM model family for several years. GLM-5.1 is their most capable release to date and the first to claim a top position on a cross-lab benchmark.
The hardware detail that matters
The training hardware is worth noting. GLM-5.1 was trained entirely on Huawei Ascend 910B chips. No Nvidia hardware was involved at any point in training. Given ongoing US export restrictions on advanced GPU chips to China, this is not just a technical footnote. It's a direct challenge to the assumption that those restrictions would constrain China's frontier AI development. Z.ai's results suggest that assumption may have been optimistic, or at minimum that Huawei's domestic AI hardware has reached a level of capability that wasn't widely anticipated.
What to be cautious about
The SWE-Bench Pro scores are self-reported by Z.ai as of early April 2026. No independent third-party lab has published corroborating results. SWE-Bench also measures a specific kind of performance: automated software engineering on real GitHub issues. It says nothing about reasoning, instruction following, creative tasks, or multimodal capabilities. GLM-5.1 is text-only, with no image, audio, or video support.
The benchmark lead is also slim. A gap of 0.7 points separates GLM-5.1 from GPT-5.4. That's within the range that methodological differences or test-set choices could influence. The result is notable, but it shouldn't be read as a definitive statement about which model is better across real-world tasks.
Why it matters regardless
The combination of MIT licensing, open weights, competitive benchmark performance, and Huawei-only training is unusual enough to pay attention to. For developers who want to run a frontier-class model locally, on air-gapped infrastructure, or without API rate limits and costs, GLM-5.1 is worth evaluating. For policymakers watching the AI hardware export control situation, it's a data point about what Chinese labs are producing without access to US silicon.
We covered the narrowing US-China AI gap in Stanford's 2026 AI Index last week. GLM-5.1's coding benchmark performance fits that broader pattern. VentureBeat and Constellation Research have additional technical coverage.
Sources
- i. huggingface.co
- ii. venturebeat.com
- iii. www.constellationr.com
- iv. medium.com
- v. www.buildfastwithai.com
Commentarii · 0