Every buying conversation in healthcare AI eventually reaches the same sentence: "It is FDA cleared." In the room, that sentence usually ends the evidence discussion. It should start it.

There are now well over a thousand AI and machine-learning enabled devices on the FDA's public list of authorized products, and the overwhelming majority reached the market through the 510(k) pathway. That pathway does not ask whether a device improves outcomes. It asks whether the device is substantially equivalent to something already on the market. Those are different questions, and confusing them is the single most expensive mistake I see health systems make.

What each pathway actually certifies

510(k) clearance is a comparison claim. The manufacturer identifies a predicate device, shows that the new product has the same intended use and similar technological characteristics, and supports that with bench testing plus, very often, a retrospective reader study on archived images. No prospective trial is required. No outcome is required. A device can be cleared without a single patient ever having been managed using its output.

De Novo authorization is what happens when there is no predicate. The FDA creates a new classification and sets special controls. That usually means more scrutiny than 510(k), and it is a reasonable positive signal, but it still does not guarantee outcome data.

PMA approval is the real bar, reserved largely for Class III devices, and it does require clinical evidence of safety and effectiveness. Very little clinical AI goes through it.

So when a vendor says "cleared," the correct follow-up question is: cleared against which predicate, and on what data? The answer is public. The clearance summary tells you the intended use statement, the study design, the sample size, and often the sites the data came from. Read it before the second meeting.

The three gaps a clearance letter will not close

The population gap. A model validated on retrospective images from three academic centers is not validated for your community hospital with different scanners, different acquisition protocols and a different patient mix. Performance drift across sites is the most consistent finding in the deployment literature, and it is almost never a software defect. It is a distribution problem.

The workflow gap. Standalone accuracy is measured with the model alone. Care is delivered by a clinician plus the model. Those systems behave differently. A high-sensitivity tool that fires too often produces alarm fatigue, and a tool that quietly agrees with the reader can produce automation bias, which is the reader deferring to a machine that happens to be wrong. Neither shows up in a reader study.

The maintenance gap. The intended use statement is frozen at clearance. The model is not. Under the predetermined change control plan framework, manufacturers can specify in advance the modifications they intend to make and how they will validate them. That is progress, and it also means the thing running in your hospital next year may not be the thing described in the letter you read this year.

How to read the regulatory record like a buyer

I use five questions, in order, and I do not move past one until it is answered in writing.

  1. What is the exact intended use statement, word for word, and does your planned deployment sit inside it? Most "off label" AI use in hospitals is not deliberate, it is drift.
  2. What was the validation data: how many patients, how many sites, what scanner or EHR vendors, what prevalence of the target condition?
  3. What is the performance at your prevalence, not theirs? Sensitivity and specificity travel between sites. Positive predictive value does not. A tool with 90 percent specificity in a low prevalence screening population will bury your team in false positives.
  4. What is the monitoring plan after go live, who owns it internally, and what threshold triggers a pause?
  5. Is there any prospective evidence at all, published or in progress, and if not, will the vendor co-author it with you?

That last question is the tell. Vendors who intend to build durable clinical products want the prospective study. Vendors selling a demo do not.

Why this is not an argument against AI

None of this means cleared AI is unsafe or unhelpful. Some of the best-supported products in the sector, in stroke triage, in diabetic retinopathy screening, in arrhythmia detection, have real prospective data and real outcome evidence behind them, and they earned it after clearance rather than before it. The point is that the clearance is the beginning of the evidence pathway, not the end of it.

The mature institutions I speak to have stopped treating regulatory status as procurement's job and started treating it as a clinical governance question, with the same committee structure they would apply to a new drug formulary. That single organisational move separates the systems that get value from clinical AI from the systems that get a shelf of expensive pilots.

Regulation tells you a device may be sold. Evidence tells you it should be used. Only one of those is your responsibility.