An AI research tool has helped doctors put names to 18 rare diseases in children whose cases had gone unsolved for years, according to a study published this week in NEJM AI by researchers at Boston Children's Hospital, Harvard Medical School and OpenAI.

The team revisited 376 de-identified case files of children with suspected genetic disorders that had resisted diagnosis, in some instances after a decade of testing and specialist review. They ran each case through OpenAI's o3 model in a "deep research" mode that reads the medical literature, weighs candidate explanations and assembles its reasoning into a report a clinician can check. Doctors then worked through the model's suggestions, ordered confirmatory tests and validated what came back. Eighteen of those cases ended with a confirmed diagnosis.

The machine did not diagnose anyone

That distinction matters, and the researchers are blunt about it. The model diagnosed no patient. It generated leads. Every diagnosis was made by qualified clinicians through the ordinary work of review, testing and confirmation, with the AI acting as a fast, tireless research assistant rather than a decision-maker. Some headlines have framed the result as software outdoing specialists. The study describes something more modest and arguably more useful: a tool that surfaces possibilities a busy human might not have time to chase down.

Rare genetic diseases are a good fit for that kind of help. There are thousands of them, the relevant findings are scattered across a sprawling literature, and one patient's symptoms can point in dozens of directions at once. A system that can read everything and never tires is well suited to the search, even when it cannot be trusted to make the final call. It is the same reason we cautioned readers, in an earlier piece, against the claim that AI already out-diagnoses your doctor.

Part of a wider health push

The study lands alongside a broader OpenAI move into medicine. The company says more than 230 million people now ask ChatGPT health and wellness questions every week, and it has been tuning its models to handle those questions more carefully, with better recognition of urgent situations and more willingness to admit uncertainty. It also recently expanded GPT-Rosalind, its model aimed at drug discovery and genomics.

What this week's result does not settle is how the tool performs outside a curated research setting, on messy live cases under time pressure, with all the liability real medicine carries. Eighteen answered cases out of 376 is a meaningful hit rate for problems that had defeated years of effort. It is not a finished product, and the researchers do not pretend otherwise.

Sources

  1. i. openai.com
  2. ii. www.techtimes.com
  3. iii. www.theneurondaily.com

Commentarii · 0

Add · a · Comment