Data Understanding & Preparation

Class Imbalance Lab

See why high accuracy can hide weak minority-class detection, and why resampling belongs inside the training workflow rather than on the evaluation set.

Lab concept guide

What to observe while you experiment

Class imbalance changes what “good performance” means because a majority-class prediction can achieve high accuracy while missing the minority class almost entirely. The lab separates prevalence, decision threshold and resampling so their effects can be inspected independently.

MechanismChange class prevalence and inspect the confusion matrix, precision, recall and PR evidence before and after any train-side resampling.
Failure modeResampling the validation/test set or celebrating accuracy while minority recall/precision collapses.
VerificationReconstruct accuracy, precision and recall from TP/TN/FP/FN and confirm that resampling changes only the training workflow, not the evaluation population.
Experiment deliberately
Create a highly imbalanced case, predict what a majority-only classifier would score, then compare accuracy with recall and PR-oriented evidence.
Evaluation set remains untouched. Only the training rows are resampled. The test prevalence therefore reflects the original problem.
Generating imbalanced classification problem…

Training geometry

Precision–recall curve

Confusion matrix & interpretation