One of the most durable predictions about modern AI is that family doctors are about to be automated. A study published in JAMA Network Open on April 14 takes a careful look at that claim, and the answer it lands on is not flattering. Across 21 frontier large language models, including the latest Claude, DeepSeek, Gemini, GPT, and Grok releases, every model failed to produce an appropriate differential diagnosis more than 80 percent of the time. Some failed in 100 percent of cases.
The work was led by researchers at Mass General Brigham, who tested the models on 29 clinical vignettes (16,254 model responses in total). Each vignette was delivered piece by piece. The models first saw the patient's age, sex, and presenting symptoms. Examination findings and lab results came in only after the model had committed to its initial reasoning. That structure mirrors how a clinician thinks, narrowing possibilities while new information arrives.
Where the models did well, and where they fell over
Final-diagnosis accuracy looked impressive: more than 90 percent for the best performers when the model was handed the full clinical picture. Differential diagnosis is harder, and that is where the models broke. Building a list of plausible explanations and revising it under uncertainty is the core of clinical reasoning, and it is also where errors in real practice are most dangerous.
The PrIME-LLM scores ranged from 0.64 for Gemini 1.5 Flash to 0.78 for Grok 4. Reasoning-optimized models did better than non-reasoning models. None reached the level the authors describe as safe for unsupervised clinical-grade deployment.
Why the myth keeps coming back
The "AI doctor" headline survives partly because benchmark numbers from chatbot vendors keep beating the United States Medical Licensing Examination. Multiple-choice questions on a known curriculum are not the same task as a tired patient walking in with a confused history. The JAMA study tests something closer to that second task, and the gap shows.
It is also worth noting what the study does not say. AI in healthcare is not useless. Radiology and pathology workflows lean on machine vision tools that are demonstrably accurate, and a separate Euronews report notes the same finding alongside high hospital adoption rates. The point is narrower: off-the-shelf chatbots, on their own, are not ready to act as primary diagnosticians.
For patients, the honest reading is that AI in the clinic still needs a human in the loop. For technologists, it is a reminder that exam scores and bedside performance measure different things. Earlier this year a separate strand of research looked at why "AI beats humans on a benchmark" headlines so often mislead. The clinical reasoning paper is a particularly clean example.
Sources: JAMA Network Open, Mass General Brigham, Euronews, News-Medical.
Commentarii · 0