4 · Data Cleaning & Missing Data

Data Preprocessing

Preprocessing converts raw variables into a representation that models can learn from without accidentally exposing validation/test information. The topic is split into focused lessons so definitions, implementation decisions, diagnostics and common failure modes can be learned separately.

How to use this topic

Learn the mechanism one decision at a time

Work through the lessons in order if the topic is new. If you already know the basics, open the specific leaf lesson that matches the operation, diagnostic or failure mode you need.

1Definition→
2Mechanism→
3Example→
4Diagnostic→
5Decision
01
Missing valuesImpute using statistics or learned models fitted on training data only. Missingness indicators can be useful when absence itself is informative. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
02
ScalingStandardisation centers/scales variance; Min–Max maps ranges; Robust scaling relies on quantiles and is less sensitive to extremes. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
03
Categorical encodingOne-hot encoding is transparent for nominal variables. Ordinal encoding requires real order. Target encoding must be cross-fitted to avoid leakage. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
04
TransformationsLog, Box–Cox/Yeo–Johnson and rank transformations can stabilise skew or variance but change interpretation. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
05
PipelinesPackage preprocessing and modelling into one pipeline so each validation fold learns transformations only from its training subset. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.