Four Clinical AI Diagnostic Tasks and How To Evaluate Them
Triage, detection, characterization, and autonomous diagnosis make different clinical claims. Here is the evidence question that belongs to each one.
AI diagnostics covers four distinct clinical tasks: triage, detection, characterization, and narrow autonomous diagnosis.
AI diagnostics is software that contributes to finding, prioritizing, or characterizing a possible disease. It is an umbrella term, not one clinical claim. Most diagnostic AI performs one of four jobs: triage, detection, characterization, or a narrowly authorized autonomous result.
Triage changes reading order. Detection marks a possible finding. Characterization estimates what a finding is. Autonomous AI produces a result without a specialist reading the underlying study. Each task requires different evidence and creates a different failure mode.
A triage alert may shorten time to review while doing nothing to improve image interpretation. A detection mark can help a reader notice a location while also creating automation bias around unmarked regions. A characterization score sounds more like an answer and therefore needs a stronger reference standard. Autonomous diagnosis carries the highest workflow burden because no specialist interpretation sits between the model and the result.
The companion guide to the four clinical AI diagnostic tasks explains how to evaluate each task without treating them as interchangeable.
Start with the intended use, the shipped threshold, the study population, and the reference standard. Sensitivity without specificity is half a result. AUROC summarizes performance across thresholds but does not tell you how the deployed threshold behaves. In screening, positive predictive value changes with disease prevalence, so a strong study result can create mostly false positives in a lower-prevalence population.
This page is the broad guide to AI diagnostics across specialties. The FDA device directory holds individual device records, while specialty and article pages explain a specific workflow, evidence question, or clinical setting.
Triage, detection, characterization, and autonomous diagnosis make different clinical claims. Here is the evidence question that belongs to each one.
Artificial intelligence in medicine uses computational models to support defined clinical, administrative, and research tasks. It does not describe one technology or one level of autonomy.