Imbalanced Learning
Imbalanced learning addresses rare classes through evaluation, sampling, weighting, thresholding and suitable objectives rather than simply maximising overall accuracy. The topic is split into focused lessons so definitions, implementation decisions, diagnostics and common failure modes can be learned separately.
Learn the mechanism one decision at a time
Work through the lessons in order if the topic is new. If you already know the basics, open the specific leaf lesson that matches the operation, diagnostic or failure mode you need.
1Definition→
2Mechanism→
3Example→
4Diagnostic→
5Decision
Start with metricsUse class-wise recall, precision, PR-AUC, balanced accuracy and confusion matrices before changing the data. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
02Class weightingWeighted losses increase the contribution of minority examples without inventing synthetic observations. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
03ResamplingRandom under/over-sampling and methods such as SMOTE change the training distribution. Resampling must happen inside each training fold. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
04ThresholdingA good ranking model may need a different decision threshold to satisfy sensitivity, precision, workload or cost constraints. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
05Rare-event validationEnsure every fold contains enough minority cases and preserve groups/time where required; otherwise estimates become unstable or optimistic. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.