One assumption running through public debate about AI is that safety benchmarks, the standard tests developers run on frontier models, are keeping pace with capability gains. The 2026 Stanford AI Index found that assumption is not well supported by evidence.
The report found that while capability benchmarks keep improving month over month, responsible AI benchmarks covering safety, fairness, and factuality are, in the report's own words, "largely absent." The majority of frontier model developers report nothing across these dimensions.
A gap that is not subtle
On capability, the field has rich, standardized benchmarks: MMLU, BIG-Bench, HumanEval, SWE-Bench, dozens more. On safety, evaluations are inconsistent, frequently proprietary, and often not published at all. When a lab says a model "passed safety testing," there is usually no standardized rubric behind that claim that an outside party can verify.
This is not a new criticism. But the scale of non-reporting documented by Stanford's AI Index is striking. The leading labs in the world, building the most capable models, are largely self-reporting on safety. Many are not reporting anything at all.
A separate warning from 100 researchers
A parallel finding came from the International AI Safety Report, compiled by over 100 researchers across more than 30 countries and led by Turing Award winner Yoshua Bengio. The report warned that the most pressing AI risks may not come from the models themselves but from the complex systems companies build around them. An agent operating autonomously inside an enterprise network is a different risk surface than the same model in a test environment. Benchmarks designed for static model evaluation may miss the risks that emerge when models are chained together with tools and real-world permissions.
What to take from this
The practical takeaway is not that AI is secretly dangerous. It is that "we ran safety tests" is a claim worth interrogating. What tests, measuring what, compared against what standard? For most frontier models today, the public cannot answer those questions. That is a problem worth naming clearly, regardless of where you stand on AI risk broadly.
Commentarii · 0