9 · Evaluation, Metrics & Diagnostics

Classification Metrics

Classification metrics answer different questions: how often predictions are correct, how well minority positives are found, how reliable probabilities are, and how ranking quality changes with threshold. The topic is split into focused lessons so definitions, implementation decisions, diagnostics and common failure modes can be learned separately.

How to use this topic

Learn the mechanism one decision at a time

Work through the lessons in order if the topic is new. If you already know the basics, open the specific leaf lesson that matches the operation, diagnostic or failure mode you need.

1Definition→
2Mechanism→
3Example→
4Diagnostic→
5Decision
01
Confusion matrixTP, TN, FP and FN are the building blocks of threshold-based metrics. Always inspect the counts as well as derived scores. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
02
Precision, recall and F-scoresPrecision measures purity of positive predictions; recall/sensitivity measures positive coverage. F1 balances them, while Fβ can emphasise recall or precision. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
03
ROC-AUC and PR-AUCROC-AUC evaluates ranking across thresholds; PR-AUC is often more informative when positives are rare because it focuses on positive predictive quality. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
04
Probability metricsLog loss and Brier score evaluate probabilistic confidence, penalising confident wrong predictions. Calibration should be inspected when probabilities drive decisions. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
05
Balanced and correlation metricsBalanced accuracy averages class recalls. MCC and Cohen’s κ can offer robust single-number summaries in imbalanced settings, but should still be accompanied by task-specific metrics. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.