Data Preprocessing
Preprocessing converts raw variables into a representation that models can learn from without accidentally exposing validation/test information. The topic is split into focused lessons so definitions, implementation decisions, diagnostics and common failure modes can be learned separately.
Learn the mechanism one decision at a time
Work through the lessons in order if the topic is new. If you already know the basics, open the specific leaf lesson that matches the operation, diagnostic or failure mode you need.
1Definition→
2Mechanism→
3Example→
4Diagnostic→
5Decision
Missing valuesImpute using statistics or learned models fitted on training data only. Missingness indicators can be useful when absence itself is informative. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
02ScalingStandardisation centers/scales variance; Min–Max maps ranges; Robust scaling relies on quantiles and is less sensitive to extremes. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
03Categorical encodingOne-hot encoding is transparent for nominal variables. Ordinal encoding requires real order. Target encoding must be cross-fitted to avoid leakage. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
04TransformationsLog, Box–Cox/Yeo–Johnson and rank transformations can stabilise skew or variance but change interpretation. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
05PipelinesPackage preprocessing and modelling into one pipeline so each validation fold learns transformations only from its training subset. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.