Calibration & Decision Thresholds
Calibration concerns whether predicted probabilities match observed frequencies; thresholding converts scores/probabilities into operational decisions. The topic is split into focused lessons so definitions, implementation decisions, diagnostics and common failure modes can be learned separately.
Learn the mechanism one decision at a time
Work through the lessons in order if the topic is new. If you already know the basics, open the specific leaf lesson that matches the operation, diagnostic or failure mode you need.
1Definition→
2Mechanism→
3Example→
4Diagnostic→
5Decision
Threshold selectionChoose thresholds using the decision objective: sensitivity target, precision target, expected cost, workload or utility. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
02Reliability diagramsBin predicted probabilities and compare mean confidence with observed outcome frequency. Deviations from the diagonal indicate miscalibration. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
03Platt / sigmoid scalingFits a logistic mapping from raw scores to probabilities on calibration data. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
04Isotonic regressionLearns a flexible monotonic calibration mapping; it can fit complex miscalibration but needs more calibration data. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
05Temperature scalingCommon in neural networks: a single temperature rescales logits, preserving class ranking while changing confidence. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.