Glossary

Calibration

Also called: probability calibration

Whether a model's stated probabilities match observed frequencies ... when it says 30 percent, does the event happen about 30 percent of the time.

A well-calibrated model's output can be read as an actual probability. Among all cases it scored at 0.3, roughly 30 percent should turn out positive. This is what makes a risk score usable in a clinical conversation.

Where This Gets Misread

Calibration is population-specific and degrades faster than discrimination when a model moves to a new setting. A model can keep ranking patients correctly while its absolute probabilities become badly wrong, and every threshold-based decision rule built on it silently breaks. Papers report AUROC far more often than calibration curves.