Classification Metrics
Classification metrics answer different questions: how often predictions are correct, how well minority positives are found, how reliable probabilities are, and how ranking quality changes with threshold. The topic is split into focused lessons so definitions, implementation decisions, diagnostics and common failure modes can be learned separately.
Learn the mechanism one decision at a time
Work through the lessons in order if the topic is new. If you already know the basics, open the specific leaf lesson that matches the operation, diagnostic or failure mode you need.
1Definition→
2Mechanism→
3Example→
4Diagnostic→
5Decision
Confusion matrixTP, TN, FP and FN are the building blocks of threshold-based metrics. Always inspect the counts as well as derived scores. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
02Precision, recall and F-scoresPrecision measures purity of positive predictions; recall/sensitivity measures positive coverage. F1 balances them, while Fβ can emphasise recall or precision. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
03ROC-AUC and PR-AUCROC-AUC evaluates ranking across thresholds; PR-AUC is often more informative when positives are rare because it focuses on positive predictive quality. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
04Probability metricsLog loss and Brier score evaluate probabilistic confidence, penalising confident wrong predictions. Calibration should be inspected when probabilities drive decisions. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
05Balanced and correlation metricsBalanced accuracy averages class recalls. MCC and Cohen’s κ can offer robust single-number summaries in imbalanced settings, but should still be accompanied by task-specific metrics. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.