AI in Dermatology: Why The Research Outpaced The Clearances
Skin lesion classification became a computer vision benchmark with headline results. Almost none of it reached clearance, and the reason is in the training data.
Image-based AI for skin lesion classification, melanoma detection, and dermatological screening.
Dermatology is well-suited to AI because of its image-centric diagnostic process. Models trained on dermoscopy images can classify skin lesions and flag potential melanoma for review, with FDA-cleared tools now available.
Dermatology has more published AI research than almost any specialty and remarkably few authorized devices, and the gap is instructive.
Skin lesion classification from photographs became a benchmark problem in computer vision, with widely reported results matching dermatologist performance on curated image sets. Very little of that reached clearance, because curated image sets are not clinical practice.
Public dermatology image datasets substantially under-represent darker skin tones. Models trained on them perform worse on the patients for whom diagnostic delay in melanoma already carries the worst outcomes. This is the clearest example in medicine of algorithmic bias arising directly from dataset composition rather than from anything subtle.
The authorized devices took a different route entirely: rather than classifying a photograph, they measure a physical property of the tissue. That sidesteps the image dataset problem, and it is not a coincidence.
Authorized devices are indicated for lesions a clinician has already judged suspicious, which means they support a referral or biopsy decision and do not address the lesion nobody examined. High sensitivity with limited specificity makes them far better at supporting a decision to biopsy than a decision not to. Consumer-facing skin apps are not cleared devices and should not be read as comparable.
2 records currently tracked.
Skin lesion classification became a computer vision benchmark with headline results. Almost none of it reached clearance, and the reason is in the training data.
Triage, detection, characterization, and autonomous diagnosis make different clinical claims. Here is the evidence question that belongs to each one.
Twelve questions that separate a clinical AI tool worth deploying from one that will look impressive in a demonstration and disappoint in production.