8 · Hyperparameter Optimisation & Model Selection

Model Selection & Comparison

Model comparison asks whether performance differences are stable, meaningful and worth the complexity rather than simply selecting the highest single score. The topic is split into focused lessons so definitions, implementation decisions, diagnostics and common failure modes can be learned separately.

How to use this topic

Learn the mechanism one decision at a time

Work through the lessons in order if the topic is new. If you already know the basics, open the specific leaf lesson that matches the operation, diagnostic or failure mode you need.

1Definition→
2Mechanism→
3Example→
4Diagnostic→
5Decision
01
Baselines firstCompare against simple baselines and current operational rules before complex models. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
02
Paired evaluationUse the same folds and test cases for candidate models so differences are paired rather than confounded by different samples. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
03
VariabilityReport fold/seed distributions, confidence intervals or bootstrap uncertainty where appropriate. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
04
Multiple objectivesCompare predictive quality with calibration, latency, memory, interpretability, fairness and maintenance cost. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
05
Final selectionFreeze the selected pipeline and evaluate once on untouched data or a prospective period before deployment. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.