For the first time in seven years, a leading AI lab has publicly admitted it built something it won't release. Anthropic announced Claude Mythos Preview on April 7, 2026, and simultaneously declined to make it available to the general public. The reason: the model's cybersecurity capabilities are, by the company's own assessment, too dangerous to hand over freely.

The benchmark numbers explain the concern. Mythos scored 93.9% on SWE-bench Verified, the standard test for automated software engineering. Claude Opus 4.6, its predecessor, scored 80.8% on the same test. On USAMO 2026, a mathematics competition designed for the country's best high school students, Mythos reached 97.6%. Opus 4.6 reached 42.3%.

The cybersecurity numbers drove the decision to restrict access. On CyberGym, a benchmark that measures the ability to find and exploit real security vulnerabilities, Mythos scored 83.1%. In practice, Claude Code running on Mythos found bugs in every major operating system and web browser tested, including some that had apparently gone undetected for decades. In controlled tests, it reproduced known vulnerabilities and produced working exploits on the first attempt 83.1% of the time.

Anthropic's response was Project Glasswing, a closed consortium of eight technology and security companies: Amazon, Apple, Cisco, CrowdStrike, Google, JPMorgan Chase, Microsoft, and Nvidia. Roughly 50 partner organizations in total have been cleared for access to Mythos Preview, strictly for defensive security work. The model is not available through the standard Claude API, on Claude.ai, or through any third-party cloud provider.

CEO Dario Amodei met with the White House on April 17 to discuss the model and its national security implications, the day after Glasswing partners received access.

What makes this notable is not just the restricted release. Most labs, when they filter a model's capabilities or restrict a deployment, do so quietly. Anthropic published what Mythos can do, explained why it's not releasing it publicly, and tied future broader access to progress on interpretability research. That's a different posture from the usual practice of simply not mentioning what gets held back.

The AI Security Institute conducted an independent evaluation and found "continued improvement in capture-the-flag challenges and significant improvement on multi-step cyber-attack simulations." In controlled conditions, the model executed multi-stage attacks on vulnerable networks and discovered exploitable vulnerabilities autonomously, replicating work that would take human security professionals days. That's the kind of finding that makes the decision to restrict access understandable, whatever you think of the precedent it sets.

Whether a broader release happens depends on Anthropic's Responsible Scaling Policy, which links deployment decisions to concrete safety evaluations. The implication is that Mythos could reach general availability once the company develops better tools for understanding what the model is actually doing internally. For now, most AI security researchers won't have access to the tool that may be the most capable at finding the vulnerabilities they study. Full details on Project Glasswing are at anthropic.com/glasswing.

Sources

  1. i. red.anthropic.com
  2. ii. www.anthropic.com
  3. iii. www.aisi.gov.uk
  4. iv. techcrunch.com
  5. v. www.axios.com
  6. vi. www.securityweek.com

Commentarii · 0

Add · a · Comment