Deeper walkthroughRead Create a Leakage Safe Pipeline as a mechanism, not a recipe
Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Choose the split strategy first. Decide whether observations may be randomly split or whether time, groups, users, patients or repeated measurements require a constrained splitter. Stage 2: Identify all learned transformations. Imputation, scaling, encoding, feature selection and PCA learn parameters from data and therefore belong inside the pipeline. Stage 3: Construct column-specific preprocessing. A ColumnTransformer can send numeric and categorical variables through different branches. Final checkpoint: Fit the chosen pipeline on development data only, then evaluate the untouched test set. The final test set remains outside model selection and tuning.
MechanismFollow the transformation
Choose the split strategy first. Decide whether observations may be randomly split or whether time, groups, users, patients or repeated measurements require a constrained splitter.
Identify all learned transformations. Imputation, scaling, encoding, feature selection and PCA learn parameters from data and therefore belong inside the pipeline.
Construct column-specific preprocessing. A ColumnTransformer can send numeric and categorical variables through different branches.
EvidenceKnow what would convince you
Assume a table has numeric columns age and income , categorical column region , and a binary target. We want training-fold medians, training-fold scaling and training-fold category mappings.
- For every learned preprocessing object, ask: “Which rows were visible when this parameter was estimated?”
- Run cross-validation on the complete pipeline rather than on preprocessed arrays.