Yes, in the one clinical area with real prospective evidence, mammography screening, AI-assisted reading measurably improves cancer detection without materially increasing recall rates. That is a genuinely strong result. It is also the exception rather than the rule across radiology AI as a category, and the gap between "cleared" and "proven" remains wide.
Where the evidence is actually strong: mammography
The AI-STREAM prospective multicenter cohort study compared breast radiologists reading mammograms with and without AI-based computer-aided detection in a real single-read screening setting, not a retrospective reader study. A separate paired, noninferiority trial published in Nature Medicine tested AI-based triage and decision support across mammography and digital tomosynthesis for breast cancer screening. Together, this is the kind of evidence base regulators and clinicians should be demanding across the category: prospective, paired, measuring the human-plus-machine system rather than the algorithm in isolation.
Where the evidence is thin: almost everywhere else
The FDA's device clearance database tells a different story. Recent 510(k) clearances for AI imaging tools, covering lung nodule detection, stroke triage on CT, and lesion characterization software, are approved on the strength of "substantial equivalence," bench testing and retrospective reader studies against archived images. Devices like RevealAI-Lung and the Median Technologies eyonis LCS platform cleared in early 2026 followed this pathway, which is standard practice and not a defect, but it means the clearance letter alone tells you almost nothing about performance in your reading room, on your scanners, at your disease prevalence.
The three numbers radiology leaders should ask for before buying
- Sensitivity and specificity at your local disease prevalence, not the vendor's validation cohort prevalence. A tool with excellent specificity in a high-prevalence academic cohort can generate a flood of false positives in a lower-prevalence community screening population, because positive predictive value moves with prevalence even when sensitivity and specificity do not.
- Radiologist time per study, measured, not estimated. AI triage is often sold on workflow efficiency. Few vendors can show measured read-time data from a comparable clinical environment rather than a lab setting.
- Prospective real-world evidence, if it exists, and if not, a commitment to build it with you. Devices like JLK-NCCT now list real-world evidence data directly in their FDA submission record, a newer and welcome trend that should become the expectation rather than the exception.
What AI reliably helps with today
- Triage and worklist prioritization. Flagging a likely acute finding, such as a large-vessel occlusion on CT, so it is read sooner rather than later in the queue. This is a timing intervention, not strictly an accuracy intervention, and it has some of the most consistent supporting data in the field.
- A second read in screening mammography, where the AI-STREAM and Nature Medicine trial data are the strongest in radiology.
- Quantitative measurement tasks, like lesion sizing or volume tracking over serial scans, where the model is doing arithmetic a human would also do but faster and more consistently.
Where the evidence does not yet support broad claims
- Standalone diagnostic replacement of a radiologist read, in any modality.
- Population-level cancer detection improvement outside screening mammography, where the trial infrastructure has simply not caught up.
- Generalization across scanner vendors and acquisition protocols without local revalidation, which the deployment literature consistently flags as the most common source of real-world performance drift.
How to evaluate a vendor pitch
Ask for the study design behind every accuracy number in the deck. Retrospective reader study on archived images is a much weaker claim than prospective, paired, real-world evidence, and the difference is not academic, it is the difference between a number that describes a lab and a number that describes your hospital.
The takeaway
Radiology AI has produced one of healthcare AI's genuinely well-supported success stories, and it took a large, expensive, prospective, multi-site trial infrastructure to get there. That is the honest cost of real evidence. Any vendor promising the same confidence level in a category that has not run those trials yet is selling you the mammography result and delivering something else.






