If it’s cleared by the FDA, it must work. Right? That’s an assumption that most clinicians make. And to be honest, I have made that assumption in the past as well. But a recent study by Abulibdeh et al. showed that only 3 of the 1,357 AI medical devices cleared by the FDA were actually tested on patient outcomes.
The study
The study, published August 19th, 2026, in PLOS Digital Health, was a systematic analysis of all AI or machine learning-enabled medical devices through December 5, 2025. The FDA database was examined and cross-referenced with ClinicalTrials.gov and PubMed.
According to the authors, only 34 (2.5%) of the 1,357 devices were linked to prospective trials. 12 (0.9%) were associated with peer-reviewed publications. And most importantly, only 3 (0.2%) examined mortality, morbidity, or readmissions (patient-centered outcomes).
The figure above, from the original publication, tells the entire story.
Accuracy is not outcomes
You might be thinking, if they didn’t study clinical outcomes, what did they study?
82% of the devices used diagnostic accuracy as their primary outcome. But accuracy as an outcome can be problematic. A device can have a 95% sensitivity and still make no difference in mortality, hospitalizations, or any patient-relevant factors. Although accuracy is important, clinical outcomes matter more.
Looking at the article findings, even the devices that underwent examination for clinical outcomes had study problems.
Pregnant patients were excluded in 42% of cardiac trials and 33% of radiology trials.
Patients who did not speak English, had cognitive impairment, and pediatric patients were all excluded.
Only 27% reported any kind of subgroup analysis. Only 3 trials examined the effects of race or ethnicity. None examined language.
75% had fewer than 500 patients.
The implications are significant for anyone practicing clinical medicine. They are also strikingly familiar.
Look at Epic’s Sepsis Model, now deployed in over half of the hospitals in the U.S. It was widely praised as evidence that AI early warning systems save lives. However, A University of Michigan study of nearly 28,000 patients found the model missed two-thirds of all sepsis cases despite generating alerts on 18% of all hospitalized patients. The model's area under the curve was 0.63, significantly worse than Epic's own published performance claims.
Epic’s tool is FDA-cleared and in your hospital already. It is a good example of how far studying accuracy with clinical outcomes gets you. An AI-based system that adds to the alert fatigue without making a real difference in patients we treat.
Why this happens: the 510(k) machine
Almost all of these AI or machine learning devices use the FDA’s 510(k) pathway for clearance. The problem is that the pathway doesn’t ask if a device works, but instead asks if a device is “substantially equivalent” to something already on the market. If an AI sepsis detector is similar to another company’s sepsis detector, it is cleared. The comparison or “predicate device” may have also been cleared by comparing to another device. So we may be able to trace a chain backwards through a series of devices that were never rigorously validated, each one’s evidence gap becoming the foundation for the next one’s approval.
This is not how new drug approval works. A pharmaceutical company has to run Phase III trials, pre-specify endpoints, and prove safety and effectiveness in large, diverse populations before the FDA begins the approval process. The new drug application pathway is lengthy and expensive. In comparison, the 510(k) pathway is equivalent to filing paperwork.
For clinicians, the distinction is invisible. We see “FDA cleared” with no stipulation regarding the methods or process. I made this argument in a piece earlier this year: regulatory status describes what the FDA did, not what the device does. The 1,357 devices in this study are evidence of that same gap. The errors and bias present in a “substantive equivalent” are propagated with no end in sight.
The specialty skew
Interestingly, there is a skew in the medical specialties targeted by these devices. Of the 1,357 FDA-cleared devices:
1059 are for Radiology. Chest X-ray readers, mammography screeners, lung nodule detectors. Only 0.3% of radiology devices have prospective trials.
126 are for Cardiology, with 9.5% having prospective trials.
62 are for Neurology, with 9.7% having prospective trials.
22 are for Anesthesiology, with no prospective trials.
88 are categorized as “Other” in the study. Yet this category has the highest prospective trial rate at 14.8%.
This data is important. As the largetst medical specialty targeted by AI devices, Radiology has nearly the lowest percent of prospective trials. It’s the specialty with the most vendor excitement and health system adoption, but the least validation.
The scale of harm
Epic’s Sepsis Model is not alone. IBM’s Watson for Oncology was marketed internationally and promised to revolutionize cancer treatment decisions. But it showed agreement with oncologists only 12–33% of the time, depending on the cancer type.
An independent analysis published in JAMA Network in August 2025 found that 6% of AI devices end up recalled, and nearly half of those within the first year. These aren’t unusual products. They are devices that passed FDA review, were used in medical practice, and then caused enough harm to be taken off the market.
These aren’t edge cases or rare failures. They’re what happens when accuracy metrics pass for evidence, and when the vendor’s own validation is believed to be sufficient. And right now, these devices are in your hospital.
The vendor pitch
The next time someone walks into your department with slides about their new AI tool, you need four things before you listen:
Prospective data, not retrospective. Retrospective validation is easy. Pick a dataset, run the model backward through it, measure how well it predicts what already happened. Prospective means the model ran forward through patients it had never seen, under real clinical conditions, before anyone knew the outcome.
Patient-centered endpoints. Not accuracy. Not sensitivity and specificity. Not area under the curve (AUC). Did mortality change? Did readmissions fall? Did length of stay shorten? Did time-to-treatment improve? If the vendor can’t answer that, it’s because they tested accuracy instead of a clinical benefit.
Validation in your population. The study data came from another health system or region. Your patients are different. Your workflows are different. Ask if they’ve validated in any hospital or system that looks like yours and struggles as you do.
Subgroup performance. How does it perform on pregnant patients? Non-English speakers? Patients with renal failure? Pediatric cases? The 27% of trials that reported subgroup analysis are the exception. Most vendors will not have that data. Their devices may fail on populations that aren’t represented in the training data, which is most populations.
How to fix the problem
The authors of this study propose a concrete three-phase roadmap:
Phase 0—pre-clearance retrospective validation on diverse datasets. Instead of proprietary training data, the authors recommend validation with diverse, publicly available or independently audited datasets. The vendor has to show the model works on populations beyond the hospital where it was built. This happens before FDA review.
Phase 1—peri-clearance prospective studies of >500 patients, embedded in real workflows. The authors recommend prospective trials with at least 500 patients, integrated into how clinicians actually work. This is randomized quality improvement: half your users see alerts, half don’t, and you measure what changes. This happens around the time of clearance.
Phase 2—post-clearance multi-center trials with >2,000 patients, with patient-centered endpoints and subgroup analyses. Two thousand patients across multiple health systems, measuring outcomes that matter to patients, with planned analyses for the populations most likely to fail. This is the NDA standard for drugs, applied to medical AI. It happens within a defined period after deployment.
Beyond the three phases, the authors also recommend mandatory prospective trial registration on ClinicalTrials.gov before the study starts (not after, which is what usually happens). CMS and private insurers should tie reimbursement to demonstrated clinical benefit, not FDA clearance, making the evidence question into a payment question.
They seem extreme, but these are standard practices for every other high-risk intervention; now applied to medical AI. They would close the gap between regulatory status and clinical validation.
Three questions worth watching
When the next sepsis model or radiology AI gets a 510(k) clearance announcement, ask yourself three things:
Should “FDA-cleared” carry a plain-English evidence disclaimer? A sentence for clinicians: “This device meets FDA standards for substantial equivalence to existing devices. It has not been independently validated on patient outcomes.” The label would change how people read clearance announcements.
At what point does accuracy-only evidence stop being enough to justify a purchase? Your health system spent millions on EMR integration and training. Someone’s making a business case to your CMO or CIO. At what threshold of accuracy do you say that’s not enough? 95%? 99%? Better to decide that now, before the pitch meeting.
Can a three-phase evidence standard survive startup funding cycles? A company with 18 months of runway can’t run a five-year Phase 2 trial. A venture-backed AI startup can’t wait for CMS reimbursement tied to outcomes. The evidence roadmap makes sense, but can the market support it?



