Every few months a headline announces that artificial intelligence has beaten doctors at their own job. The studies behind those headlines are real, peer-reviewed, and often impressive. The leap many readers make next, that a chatbot can now stand in for a physician, is where the evidence stops cooperating.
Take the study that drew the loudest coverage. Researchers at Harvard Medical School and Beth Israel Deaconess fed real patient records to one of OpenAI's reasoning models and found it included the correct diagnosis in its answer roughly 80% of the time, outscoring two experienced physicians on the same cases. That result is genuine and worth taking seriously. It is also narrower than it sounds. The doctors were working from limited written information, the same documentation the model saw, not a live patient they could examine, question, or send for tests.
That distinction is the heart of the matter. Diagnosis on paper is not the same task as diagnosis in a room. The American College of Emergency Physicians noted that the model was measured against internal medicine doctors, while emergency specialists trained for exactly this kind of rapid assessment would likely have done better. One of the study's own authors put it plainly: the findings do not mean AI replaces doctors.
What the benchmark leaves out
A clinician does far more than match symptoms to a likely cause. They notice the patient who downplays their pain, decide which test is worth the cost and the risk, weigh a treatment against a person's other conditions, and carry the responsibility when a judgment call goes wrong. Today's models still stumble in ways that matter at the bedside. They state false facts with total confidence, inherit biases from their training data, and lack the everyday common sense that tells a human when something simply does not add up.
This is why the most interesting research has stopped framing it as a contest. Studies of human-AI collectives, where the model and the clinician each weigh in, tend to beat either one working alone. A recent Stanford and Harvard review of clinical AI reached a similar conclusion about what holds up in practice: these systems are powerful aids to a doctor's reasoning, not substitutes for it. The model widens the list of possibilities and catches what a tired human misses. The human decides.
So is AI better than your doctor? On a multiple-choice version of the question, sometimes, yes. On the actual job, no, and the people running the studies are the first to say so. The useful question is not who wins. It is how a good clinician and a good model work together, because that pairing already outperforms both. The myth is not that AI is good at diagnosis. It is that diagnosis is the whole of medicine.
Commentarii · 0