Researchers at Carnegie Mellon University have built a hospital that does not exist. Its patients are invented, its charts are synthetic, and yet, in a blind test, physicians could barely tell its records apart from real ones. The project, called Synthetic Hospital, is an attempt to solve a stubborn problem in medical AI: how do you test these systems fairly when the data you need is locked away for privacy reasons?
The benchmark holds 1,268 patients followed over time and 5,602 clinical encounters. Every diagnosis, finding, and time relationship is grounded in standard medical vocabularies such as ICD-10-CM, SNOMED CT, and LOINC, with a documented trail back to the source material each detail came from. That grounding matters. It gives researchers a verifiable answer key, something real patient records rarely provide because the full truth of a case is often scattered or unknown.
Realistic enough to fool the experts
The realism is the headline finding. In a blinded review, doctors distinguished the synthetic charts from genuine ones only about 53 percent of the time, close to a coin flip. That is the point. If the records were obviously fake, no test built on them would tell you much about how a model handles the real thing.
The results for the AI were more sobering. Across ten frontier and open models, the best system rebuilt a patient's list of problems with a severity-weighted score of 0.73, level with the average of seven physicians on a matched set of cases. But it trailed the best doctor, who scored 0.89. More striking, when asked to summarize a chart, the models missed roughly half of the clinically relevant findings. A summary that drops half of what matters is not a small flaw in a hospital.
A useful dose of reality
The study reads as a corrective to the more excited claims about AI in medicine. These tools can match the average clinician on some structured tasks, which is genuinely useful, and they fall short of the best and lose important detail when the job is open-ended. That gap is exactly what a good benchmark is meant to expose.
It also fits a run of recent stories about AI meeting the friction of real healthcare, from insurers blaming AI billing tools for nearly a billion dollars in added costs to doctors drawing a line at AI beyond the scan. A fake hospital, oddly enough, may give the field one of its more honest measuring sticks.
Sources
- i. arxiv.org
- ii. hermes-ai.net
- iii. www.theneuron.ai
Commentarii · 0