You have heard some version of this one. AI can already read a scan better than a radiologist, radiology is a solved problem, medical schools should stop training people for a job that will not exist. Geoffrey Hinton said something close to it in 2016, suggesting we stop training radiologists altogether. Ten years on there is a global shortage of radiologists and no shortage of headlines repeating the claim.
The claim is not invented from nothing. Narrow imaging models genuinely do beat human performance on specific, well-defined tasks: flagging a particular fracture pattern, screening for diabetic retinopathy, triaging a stroke study. Those results are real and clinically useful. The trouble starts when they get generalised into a statement about diagnosis as a whole.
What the hardest benchmark found
Radiology's Last Exam, published as RadLE, set out to test frontier multimodal models against human experts on genuinely difficult diagnostic cases. The results were not close.
Board-certified radiologists scored 83 per cent. Radiology trainees, still in training, scored 45 per cent. The models: GPT-5 at 30 per cent, Gemini 2.5 Pro at 29, OpenAI's o3 at 23, Grok-4 at 12, and Claude Opus 4.1 at 1 per cent. Every model tested came in significantly below the trainee cohort, and nowhere near the consultants.
Two caveats belong here, in fairness. These were deliberately hard cases, not routine studies, so the numbers are a ceiling test rather than a picture of everyday workload. And the models tested are now a generation old, which in this field means something. A rerun with current frontier models would produce better numbers. It would need to produce dramatically better ones to change the conclusion.
The failure modes are the interesting part
The researchers catalogued how the models got things wrong, and the taxonomy is more informative than the scores. Perceptual errors included missing findings that were plainly visible, hallucinating pathology that was not there, and correctly spotting something but placing it in the wrong anatomy. Interpretive errors included attributing findings to the wrong diagnosis and settling on an answer early without considering alternatives. There were also communication errors, where a model's own reasoning contradicted the summary it produced.
That last category should give anyone pause. A system whose stated conclusion disagrees with its own working is difficult to supervise, because the part a clinician reads is not reliably the part that did the thinking.
Confidence is the actual problem
The follow-up benchmark, RadLE 2.0, changes what is being measured. Rather than scoring accuracy alone, it asks whether a model's confidence is justified and whether it knows when to hand over to a human. Correct answers earn confidence-weighted credit, wrong answers take a confidence-weighted penalty, and an honest "I don't know" scores neither.
A confidently wrong diagnosis is far more dangerous than an admitted uncertainty, and most models are poor at the latter.
This is the right thing to measure. In a working radiology department the model is not the final word, a human is. The question that matters is whether the model's uncertainty is legible enough for that human to know when to look harder. A system that is 70 per cent accurate and reliably flags its own shaky cases is more useful than one that is 80 per cent accurate and sounds equally certain throughout.
Under that scoring, a model that guesses confidently and often will fall down the rankings even where raw accuracy looks competitive. Which is a reasonable description of what current systems do.
So where does that leave the claim
Not supported, in the strong form. Supported in a narrow form that is much less exciting than the headline. AI reads specific findings very well, drafts reports competently, triages worklists, and saves radiologists real time. It does not do open-ended diagnosis at consultant level, and on the hardest available test it does not reach the level of someone still in training.
The shortage of radiologists is genuine, and there is a decent case that these tools help stretch the workforce that exists. That case does not require the claim that the tools have replaced anyone. It is the same pattern we found looking at whether AI designs drugs on its own, and at what benchmark scores actually tell you. A real capability, doing real work, described in terms that overshoot it by some distance.
If someone tells you radiology is finished, the fair question back is which benchmark they are citing, and how the humans did on the same one.
Sources
- i. arxiv.org
- ii. crashlab.in
- iii. www.emergentmind.com
- iv. www.buildfastwithai.com
Commentarii · 0