Ground Truth
Also called: reference standard, gold standard
The labels a model is trained and evaluated against, which define the ceiling on what the model can learn to do.
Ground truth is the reference the model is taught to reproduce: a radiologist's reading, a pathology result, a chart-derived outcome, or a consensus panel. Every performance figure is measured relative to it.
Where This Gets Misread
A model cannot exceed the quality of its labels, and clinical ground truth is frequently imperfect. If labels came from single-reader interpretation, the model learned that reader's errors and biases along with their skill. When a model disagrees with a clinician, that is exactly what it was trained to consider wrong ... which is worth remembering before assuming the model is the one that erred.