The benchmark numbers from AI research have become genuinely striking. Frontier models now match human scores on PhD-level science tests, topped a coding benchmark that sat at 60% just one year ago, and outperform doctors on some diagnostic tasks. So it might be surprising that when scientists put AI agents to work on real research, the best systems currently complete roughly half as much as a human expert with a PhD.
That finding comes from a Nature analysis published this spring, which examined how top AI agents performed on complex, open-ended scientific tasks rather than standardized tests or fixed benchmarks. The conclusion, drawing on data compiled for the Stanford AI Index 2026, is that the gap between benchmark scores and practical scientific capability remains wide.
Why Benchmarks and Research Are Different Things
These two facts can both be true at once because benchmarks and real research ask very different things. Tests like GPQA (graduate-level scientific knowledge) ask models to answer multiple-choice questions drawn from published material. AI systems are genuinely good at this. But actual research asks for something else: forming original hypotheses, designing experiments, navigating ambiguity, weighing conflicting evidence, and connecting findings across disciplines in ways that have not been done before.
Yolanda Gil at the University of Southern California, who helped lead the Stanford AI Index, noted that researchers have embraced AI systems as useful collaborators despite these limits. The technology works well for literature review, data analysis, and routine tasks. But an autonomous agent that independently runs a research program at expert level remains a future prospect, not a current reality.
What AI Is Actually Changing in Science
None of this means AI is not affecting scientific research. It is, and meaningfully. A Science analysis found that AI tools have accelerated scientific output, though with a possible downside: teams using AI heavily may be concentrating on faster, incremental work at the expense of higher-risk investigations. More papers does not automatically mean more discovery.
The distinction worth holding onto is between AI as a capable assistant and AI as an autonomous research agent. The first is already here and genuinely useful. The second, according to the best available evidence, is not.
Sources
- i. www.nature.com
- ii. hai.stanford.edu
- iii. www.science.org
- iv. csnsf.org
Commentarii · 0