The U.S. Food and Drug Administration has authorized more than 1,300 artificial intelligence tools for use in patient care. According to a study published on 19 August in PLOS Digital Health, only three of them have been tested to see whether they make patients any healthier.

Rawan Abulibdeh of the University of Toronto and colleagues went through 1,357 authorized devices and looked for evidence tied to patient outcomes, measures like mortality or hospital admissions rather than accuracy on a benchmark. The authors reported that they had expected the evidence base to be thin, but not this thin. As they put it, a device can clear the bar without anyone showing it helps a single person.

Cleared by comparison, not by results

The gap traces back to how most of these tools reach the market. To be authorized, an AI device usually only has to show what regulators call substantial equivalence to a product already on sale. Developers are not required to prove that the new tool improves anyone's health, or that the benefit reaches different groups of patients evenly.

That pathway made sense for simpler hardware, where a new model of an old instrument really is much like the last one. It sits less comfortably with software that reads mammograms, calculates cardiovascular risk, or helps plan surgery. These tools increasingly shape the decisions a clinician makes, and the study's point is that the shaping is happening ahead of the proof.

Why the finding matters now

None of this means the devices are harmful, and the authors do not claim they are. A tool can be accurate at the narrow task it was built for and still lack any study showing it changes what happens to the patient. The problem is that accuracy on a test set and benefit in a clinic are not the same thing, and right now the second one is rarely measured.

The finding lands in a live debate about how much weight to give an algorithm at the bedside. Research this month has already shown that the value of medical AI depends heavily on who is using it, and that a confident tool can pull an expert toward the wrong answer as easily as the right one. Add a thin outcomes record to that picture and the case for caution grows. The authors argue the fix is not to pull the tools, but to require the kind of follow-up testing that would tell doctors, and patients, whether they are actually working.

Sources

  1. i. www.news-medical.net
  2. ii. medicalxpress.com
  3. iii. www.insideprecisionmedicine.com
  4. iv. www.scimex.org

Commentarii · 0

Add · a · Comment