Here is a test. Ask an AI model to solve a problem from the International Mathematical Olympiad, one of the hardest math competitions in the world, calibrated to challenge the best high school mathematics students from across the globe. Then ask it to tell you the time from an analog clock.
According to Stanford's 2026 AI Index, Gemini Deep Think earned a gold medal at the International Mathematical Olympiad last year. The same report finds that the best available AI model reads analog clocks correctly just 50.1 percent of the time. That is barely better than guessing.
Researchers call this the jagged frontier. AI is not generally intelligent. It is extraordinary at some things and surprisingly poor at others, often in ways that do not map to any human intuition about difficulty. The problem is that the spectacular achievements get the headlines. The clock-reading failure rate becomes a footnote.
The pattern runs throughout the Stanford report. AI systems can answer PhD-level questions in biology and chemistry at rates that would challenge many graduate students. Household robots, meanwhile, complete their assigned tasks correctly only 12 percent of the time. The frontier is genuinely jagged: superhuman in some directions, barely functional in others, with no reliable way to predict which domain falls into which category.
This matters because a lot of decisions about AI deployment are made by people who have absorbed the headline version of AI capability. They hear that a system outperforms the best humans at chess, Go, and now olympiad mathematics, and they assume broad competence. That assumption is how things go wrong in practice.
The explanation for the jagged pattern is not mysterious. Large language models learn from text. Olympiad mathematics appears in training data in structured, well-documented form. Reading an analog clock requires spatial and physical reasoning that text corpora do not cover well. The model does not have a general capacity for understanding that transfers cleanly between domains. It has learned patterns that do or do not apply to whatever it is being asked to do.
None of this is a reason to discount what AI systems do well. Olympiad-level mathematical reasoning is a genuine capability, and the research record is real. But there is a meaningful difference between a tool that excels at specific, well-supported tasks and a general intelligence that can be trusted with whatever you put in front of it. Current AI systems are the former. Treating them as the latter is one of the more common and consequential mistakes in how the technology gets deployed today.
The clock problem is a useful anchor. Before handing a task to an AI system, it is worth asking whether the relevant domain looks more like olympiad mathematics or telling time.
Sources
- i. hai.stanford.edu
- ii. hai.stanford.edu
Commentarii · 0