AI in healthcare, despite substantial venture capital investment exceeding $40 billion this decade in AI biopharma alone, is encountering a significant challenge in providing concrete clinical evidence of its benefits. While AI has shown promise in identifying drug candidates and assisting with some medical decisions, it has largely failed to prove these predictions in real-world clinical settings and improve patient-centered outcomes. This gap between potential and proven impact is highlighted by a recent study revealing that less than 1% of the more than 1,300 FDA-cleared AI medical devices have been evaluated for patient outcomes like mortality or readmissions.

The majority of FDA-cleared AI devices, particularly those gaining clearance via the 510(k) pathway which only requires showing "substantial equivalence" to existing tools, lack rigorous clinical trials. The study found that only 34 (2.5%) of these devices were linked to registered prospective trials, with an even smaller fraction (0.9%) publishing results, and a mere three devices (0.2%) demonstrating patient-centered outcomes. Industry sponsorship dominates the limited evidence, with 94% of trials being industry-led and focused on accuracy metrics rather than patient benefits.

This evidence gap is particularly acute in specialties like radiology, which accounts for 78% of cleared devices but only 1% linked to prospective trials. Experts emphasize the need for evidence that measures actual patient outcomes over algorithmic accuracy, suggesting that regulatory bodies like the FDA should mandate pre-registration of prospective trials for all Class II and III AI devices. They also recommend validating devices in the populations where they will be used and linking reimbursement to demonstrated clinical benefit.

The difficulties in gathering prospective evidence for AI are manifold: conversations with AI are open-ended, context is crucial and varies widely, real-world deployment introduces unpredictable human factors, and such evaluations are time-consuming and expensive. Furthermore, AI systems are constantly updated, making evidence tied to one version quickly obsolete. These challenges underscore why trust in medical AI must be earned through rigorous, real-world studies, rather than merely through benchmarks or rapid regulatory clearance, especially as AI chatbots reportedly misdiagnose over 80% of early medical cases.