Dermatology has produced more published AI research than almost any specialty and remarkably few authorized devices. The gap is the most instructive thing about the field.
Skin lesion classification from photographs became a benchmark problem in computer vision around 2017, with widely reported results matching or exceeding dermatologist performance on curated image sets. Those results were real. Very little of that work reached clearance, because a curated image set is not clinical practice.
The training data problem is unusually severe here
Public dermatology image datasets substantially under-represent darker skin tones. A model trained on them performs worse on exactly the patients for whom diagnostic delay in melanoma already carries the worst outcomes.
This is the clearest example in medicine of algorithmic bias arising directly from dataset composition rather than from anything subtle. There is no proxy variable to untangle and no hidden confounder to discover. The images were not there.
It is also a good illustration of why external validation matters more than headline accuracy. A model reporting dermatologist-level performance on a dataset drawn from one population has demonstrated something narrower than the headline suggests.
The authorized devices went around the problem
Here is the interesting move. Rather than classifying a photograph, the devices that reached authorization measure a physical property of the tissue ... elastic scattering spectroscopy in one case, electrical impedance in another.
That sidesteps the image dataset problem entirely, because the measurement does not depend on pigmentation the way photographic classification does. It is not a coincidence that this is the route that got through.
The authorized user is the point
The 2024 De Novo authorization in this space was not cleared for dermatologists. It was authorized for primary care.
Dermatologists already assess lesions well. Most suspicious lesions, however, are first seen in primary care, where the miss rate is meaningfully higher and specialist access can be slow. A device that supports a referral decision in that setting attacks the actual gap, and the comparator is not a dermatologist ... it is whether the patient gets referred at all.
This is the same logic that made autonomous retinal screening work: value is created where the specialist is not.
What these devices do not do
They support a decision. They do not diagnose. They are indicated for lesions a clinician has already judged suspicious, which means they do nothing about the lesion nobody examined.
High sensitivity with limited specificity makes them far better at supporting a decision to biopsy than a decision not to. A reassuring result on a clinically concerning lesion does not override clinical judgment, and reading it as permission to skip a biopsy is the failure mode to guard against.
Consumer apps are a different category entirely
Consumer-facing skin analysis apps are not cleared devices, were not evaluated as devices, and should not be read as comparable to anything on the authorized device list. They carry the same dataset composition problem with none of the regulatory scrutiny, and a false reassurance delivered to a patient with no clinician involved is the worst version of this technology.
See the dermatology specialty hub for the current state of the field.