Medical artificial intelligence has become much easier to find than proof that it makes patients healthier. AI tools can already classify medical images, predict risk and suggest next steps to clinicians. The harder question, highlighted by the Financial Times, is whether those capabilities improve what happens to people in ordinary care.
A 2026 evidence census in PLOS Digital Health reviewed 1,357 AI and machine-learning medical devices authorized by the US Food and Drug Administration through Dec. 5, 2025. Only 34 were linked to registered prospective trials, 12 had posted results, 12 had peer-reviewed publications and three reported patient-centered outcomes such as mortality, morbidity or readmission.
Those numbers describe the publicly traceable evidence, not everything regulators or manufacturers may know. The FDA says devices on its list met applicable premarket requirements, while also cautioning that public summaries omit much of the information submitted in applications. The PLOS authors likewise said a lack of published outcome studies does not necessarily mean regulatory oversight was deficient.
Accuracy is not the same as better care
Many medical AI studies measure sensitivity, specificity or performance on a curated dataset. Those measures matter, but they are not the same as showing that a tool reduces complications, improves quality of life or helps patients live longer.
A separate JAMA Network Open analysis examined 903 AI-enabled devices on the FDA list through August 2024. It found that 97.1 percent had gone through the 510(k) pathway, which generally asks whether a device is substantially equivalent to an existing product. Among the 505 devices with a reported clinical performance study, only 12 studies were randomized, and the study design was missing for 259. Information about performance by age and sex was also frequently absent from public records.
That creates a practical problem for hospitals choosing between products. A strong result in one dataset or institution may not hold when patient populations, equipment or clinical workflows change. Thin subgroup reporting also makes it harder to know whether a system performs consistently across the people who will encounter it in practice.
A real-world trial shows both promise and limits