9 · Evaluation, Metrics & Diagnostics

Classification Metrics: Ranking & Probability

Classification Metrics: Ranking & Probability groups the core ideas a learner needs at the 9 · evaluation, metrics & diagnostics stage. Work through the lessons in order when new to the area, or use them independently as a reference when implementing an analysis.

How to use this topic

Learn the mechanism one decision at a time

Work through the lessons in order if the topic is new. If you already know the basics, open the specific leaf lesson that matches the operation, diagnostic or failure mode you need.

1Definition→
2Mechanism→
3Example→
4Diagnostic→
5Decision
01
ROC curveROC curve is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness depends on whether that aspect matches the real decision cost, class prevalence and intended model output.
02
ROC-AUCROC-AUC is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness depends on whether that aspect matches the real decision cost, class prevalence and intended model output.
03
Precision-recall curvePrecision-recall curve is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness depends on whether that aspect matches the real decision cost, class prevalence and intended model output.
04
Average precision / PR-AUCAverage precision / PR-AUC is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness depends on whether that aspect matches the real decision cost, class prevalence and intended model output.
05
Log lossLog loss is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness depends on whether that aspect matches the real decision cost, class prevalence and intended model output.
06
Brier scoreBrier score is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness depends on whether that aspect matches the real decision cost, class prevalence and intended model output.
07
Top-k accuracyTop-k accuracy is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness depends on whether that aspect matches the real decision cost, class prevalence and intended model output.
08
MCC and Cohen’s kappaMCC and Cohen’s kappa is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness depends on whether that aspect matches the real decision cost, class prevalence and intended model output.