Here is the finding that should shape how medical AI is deployed, and mostly does not.
Mammography computer-aided detection was widely deployed across the United States, reimbursed, and broadly assumed to be helping. Large studies later found no improvement in cancer detection and an increase in recall rates. The software was not especially inaccurate. It changed how radiologists read.
That is automation bias, and it is the mechanism by which a technically sound tool becomes a clinical liability.
Two forms, and the second one is worse
Commission errors are accepting an incorrect recommendation. The system flags something, the clinician agrees, the system was wrong. These are visible in retrospect and relatively easy to study.
Omission errors are failing to notice what the system did not flag. Nothing happened. Nothing was recorded. The finding was simply not seen, and there is no artifact anywhere indicating that a tool contributed.
Omission errors are the reason a detection tool can raise standalone accuracy and lower reader performance at the same time. An unmarked region acquires an implicit reassurance nobody designed it to carry, and the clinician's search behavior changes without their awareness.
This is not a medicine problem
Automation bias is documented across aviation, process control, maritime navigation, and military systems, in populations of highly trained professionals operating under formal procedures requiring independent verification.
Assuming physicians are exempt requires believing that clinical training confers a resistance that pilot training does not. There is no evidence for that, and the mammography result is evidence against it.
Why it undermines the regulatory model
Nearly all authorized medical AI is assistive. The clearance says the clinician remains responsible for the decision, and that framing is what keeps the evidence bar manageable.
The same logic underpins the exemption for most physician-facing tools: a human in the loop reviews the output, therefore it is not a device.
Both rest on an empirical claim about clinician behavior, and the evidence on that claim is not encouraging. Assistive describes the intended workflow. It does not describe what happens in week twelve of a deployment, when the tool has been right often enough to have earned trust it may not deserve on this particular case.
Alert fatigue is the same problem from the other end
Override rates for drug interaction alerts routinely exceed 90 percent. Clinicians dismiss them reflexively because the great majority are not clinically relevant, and dismissal becomes a motor pattern rather than a decision.
So the two failure modes are: trust the system too much, or ignore it entirely. Both are behavioral responses to a system whose specificity does not match the attention available. Neither is fixed by making the model more accurate.
Any new AI tool that generates notifications is entering an environment that is already saturated. The flag does not arrive into empty attention.
What actually helps
Design for omission errors specifically. Reading protocols where the clinician completes an independent assessment before the AI output is displayed preserve the unbiased read. This costs time, which is why it is rarely done, and it is the single most effective mitigation available.
Set specificity thresholds against attention, not against the ROC curve. A tool generating more flags than a department can meaningfully evaluate will be ignored, and its measured sensitivity becomes irrelevant.
Monitor acceptance rates as a safety metric. A rising acceptance rate over time is ambiguous ... it means the tool improved or the reviewer disengaged, and those are indistinguishable without independent audit. Sample and audit.
Be honest about what unmarked means. Training should state explicitly that an absence of flags carries no information, because clinicians will behave as though it does unless told otherwise, repeatedly.
The uncomfortable conclusion
A clinical AI tool cannot be evaluated as a standalone artifact. Its effect on outcomes depends on how it changes the behavior of the person using it, and that effect can be negative while every performance metric on the datasheet is excellent.
Which means standalone performance figures, the thing regulatory submissions are largely built from, are answering a question that is not the clinical one. Prospective evaluation in the real workflow is the only design that captures it, and it is what almost no authorized device has behind it.