Glossary
Calibration
Also called: probability calibration
Whether a model's stated probabilities match observed frequencies ... when it says 30 percent, does the event happen about 30 percent of the time.
A well-calibrated model's output can be read as an actual probability. Among all cases it scored at 0.3, roughly 30 percent should turn out positive. This is what makes a risk score usable in a clinical conversation.
Where This Gets Misread
Calibration is population-specific and degrades faster than discrimination when a model moves to a new setting. A model can keep ranking patients correctly while its absolute probabilities become badly wrong, and every threshold-based decision rule built on it silently breaks. Papers report AUROC far more often than calibration curves.