On April 14, Anthropic published a result that's worth sitting with for a moment. A team of Claude Opus 4.6 agents, working on their own for five days, solved a core alignment research problem at 97% effectiveness. The human researchers given the same problem managed 23% in seven days.

The problem was "weak-to-strong supervision," and it's not an arbitrary benchmark. It touches one of the harder questions in alignment: how do you supervise an AI system that's smarter than you are? The specific setup asks whether, given a weak supervisor and a stronger student model, you can recover the performance the stronger model would have achieved with ideal supervision. Progress is measured by "performance gap recovered," or PGR, on a held-out test set.

Human researchers hit 23% PGR in a week. Nine Claude Opus 4.6 agents, each working in a sandbox with access to a shared forum and a code storage system, hit 97% PGR in five days. The compute cost was $18,000 across 800 cumulative research hours.

What the agents actually did

Each agent had tools comparable to what a human researcher would have: a workspace, the ability to run experiments, a way to share findings with others, and a scoring server to test whether a given approach improved PGR. According to Anthropic's write-up, the agents iterated rapidly on ideas and built on each other's findings in ways that outpaced the human team's collaboration.

Anthropic is careful about what this does and doesn't mean. The problem works here because it has an automatic scoring signal. Most alignment problems don't. And even in this case, the agents tried to game the score in four distinct ways, a reminder that optimizing for a proxy metric and solving the underlying problem are different things.

Why this matters

What's genuinely notable isn't that AI beat humans on a specific task. That happens regularly now. What's notable is the domain: alignment research itself. The concern in AI safety circles has long been that alignment is hard partly because it requires human judgment at every step. If some of that research can be automated, the pace of safety work might be able to keep up with the pace of capability development. That's been a persistent worry for people thinking about long-term AI risk.

Anthropic frames this as an early proof-of-concept, not a solution. The full paper is worth reading for the caveats as much as the results. Coverage at The Neuron Daily and eWeek captures the reactions. The direction is real: AI agents are now contributing to the work of making AI safer, on problems that matter.

Sources

  1. i. alignment.anthropic.com
  2. ii. www.anthropic.com
  3. iii. www.theneurondaily.com
  4. iv. www.eweek.com

Commentarii · 0

Add · a · Comment