AI diagnostics is not one thing. It is four things with different regulatory burdens, different failure modes, and different consequences when they are wrong, and almost every confused conversation about medical AI comes from treating them as interchangeable.

The four categories are triage, detection, characterization, and autonomous diagnosis. A device authorized for one is not authorized for the others, and a vendor describing its triage tool as a diagnostic system is making a claim its clearance does not support.

Triage: changing the order, not the reading

A computer-aided triage device analyzes a study, decides it probably contains a time-critical finding, and moves it up the worklist. It does not mark the image. It does not produce a diagnosis. It does not remove anything from the queue.

The clinical claim is time, not accuracy, and the regulatory evaluation follows: these devices are assessed on sensitivity and time-to-notification. A high false-positive rate is entirely compatible with clearance.

That has a consequence worth internalizing. A triage device is allowed to be wrong often, as long as it is rarely wrong in the direction of missing something. In a department that already reads studies promptly, reordering the queue achieves nothing at all, and the benefit that justified the purchase never appears.

Detection: marking a location

Detection marks candidate findings on the image while the clinician reads. It says look here. It does not say what the thing is.

This category has the longest and most cautionary history in medical AI, and it predates deep learning entirely. Mammography computer-aided detection was widely deployed across the United States, reimbursed, and broadly assumed to be helping. Large studies later found no improvement in cancer detection and an increase in recall rates.

The mechanism behind that failure is the thing to carry forward. The software was not especially inaccurate. It changed how radiologists read, and an unmarked region acquired an implicit reassurance nobody intended it to carry. Automation bias turned a technically reasonable tool into a null result.

Characterization: saying what it is

Characterization estimates what a finding is ... benign or malignant, high or low risk. That is a substantially stronger claim and attracts a correspondingly higher evidence burden.

It is also the category most likely to be deferred to, because the output arrives as a conclusion rather than a prompt. A wrong location mark invites a second look. A wrong characterization sounds like an answer.

Autonomy: no clinician reads the study

Autonomous AI produces a clinical result with nobody interpreting the underlying data. In the United States this is currently a small category dominated by retinal screening for diabetic retinopathy.

It required a different class of evidence: prospective validation, at primary care sites, with the intended non-specialist staff, in the intended workflow. Most authorized medical AI has nothing resembling that behind it.

The reason the bar is higher is liability rather than difficulty. When no clinician read the study, responsibility for a missed finding does not sit with a reader, and the field has not settled where it does sit.

How to read a performance claim

Four questions defuse most diagnostic AI marketing.

Which category is this? The answer determines what the clearance actually permits.

What is the sensitivity and specificity at the shipped threshold? AUROC is the most quoted metric and the least actionable ... it summarizes performance across thresholds nobody would use, and two models with identical AUROC can behave completely differently at the operating point.

What was the disease prevalence in the evaluation set? This is where screening claims quietly collapse. A model with excellent sensitivity and specificity applied to a low-prevalence population produces mostly false positives, and no improvement in the model fixes it. Evaluation sets enriched with positive cases produce a positive predictive value that will not survive real prevalence.

Was it validated externally? Performance drops on external validation are the norm, and drops of ten or more points are routinely reported. A model with excellent internal validation and no external validation has been shown to work in exactly one place.

What this means in practice

The useful question about a diagnostic AI tool is never whether it is accurate. It is what task it was validated on, in whom, against what reference standard, and at what prevalence. Everything else is downstream of those four answers.

Browse the FDA device tracker to see which category each authorized device actually falls into.