Capstone Project · Lesson 82

Create a Leakage Safe Pipeline

A leakage-safe pipeline keeps every data-dependent preprocessing operation inside the training/resampling boundary.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What makes a pipeline leakage-safe

A leakage-safe pipeline places every data-dependent preprocessing operation inside the resampling boundary. Imputation values, scaling parameters, category mappings, selected features and dimensionality-reduction components are learned from training data only. The held-out data are transformed with those learned parameters but never participate in fitting them.

This is more than code organisation. It protects the validity of model evaluation. If the complete dataset is used to choose features, calculate means or standard deviations, or learn encodings before cross-validation, the validation observations influence the model-building process indirectly. The resulting metric can look better than the performance that would be achieved on genuinely unseen data.

Deeper walkthrough

Read Create a Leakage Safe Pipeline as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Choose the split strategy first. Decide whether observations may be randomly split or whether time, groups, users, patients or repeated measurements require a constrained splitter. Stage 2: Identify all learned transformations. Imputation, scaling, encoding, feature selection and PCA learn parameters from data and therefore belong inside the pipeline. Stage 3: Construct column-specific preprocessing. A ColumnTransformer can send numeric and categorical variables through different branches. Final checkpoint: Fit the chosen pipeline on development data only, then evaluate the untouched test set. The final test set remains outside model selection and tuning.

Mechanism

Follow the transformation

Choose the split strategy first. Decide whether observations may be randomly split or whether time, groups, users, patients or repeated measurements require a constrained splitter.

Identify all learned transformations. Imputation, scaling, encoding, feature selection and PCA learn parameters from data and therefore belong inside the pipeline.

Construct column-specific preprocessing. A ColumnTransformer can send numeric and categorical variables through different branches.

Evidence

Know what would convince you

Assume a table has numeric columns age and income , categorical column region , and a binary target. We want training-fold medians, training-fold scaling and training-fold category mappings.

  • For every learned preprocessing object, ask: “Which rows were visible when this parameter was estimated?”
  • Run cross-validation on the complete pipeline rather than on preprocessed arrays.
Useful distinctionSafe: Split first; fit learned preprocessing and feature selection inside the training fold; transform held-out data with the fitted objects.
Visual demonstration of Create a Leakage Safe Pipeline
Visual demonstration: use the diagram to trace the main objects and state changes involved in Create a Leakage Safe Pipeline.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Choose the split strategy first

Choose the split strategy first. Decide whether observations may be randomly split or whether time, groups, users, patients or repeated measurements require a constrained splitter.

Transformation focus: keep the input and produced parameters/result separate so the change is observable and reproducible.
How it works

Keep learned preprocessing inside each fold

  1. Choose the split strategy first. Decide whether observations may be randomly split or whether time, groups, users, patients or repeated measurements require a constrained splitter.
  2. Identify all learned transformations. Imputation, scaling, encoding, feature selection and PCA learn parameters from data and therefore belong inside the pipeline.
  3. Construct column-specific preprocessing. A ColumnTransformer can send numeric and categorical variables through different branches.
  4. Join preprocessing and estimator. A top-level Pipeline makes the complete transformation-and-model chain a single estimator.
  5. Cross-validate the complete object. Each fold receives a newly fitted preprocessor and model, preventing held-out observations from affecting learned preprocessing state.
  6. Fit the chosen pipeline on development data only, then evaluate the untouched test set. The final test set remains outside model selection and tuning.
Worked demonstration

Numeric and categorical preprocessing without leakage

Assume a table has numeric columns age and income, categorical column region, and a binary target. We want training-fold medians, training-fold scaling and training-fold category mappings.

Demonstration

ColumnTransformer inside Pipeline

# Step 1 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.compose import ColumnTransformer
# Step 2 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.impute import SimpleImputer
# Step 3 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.linear_model import LogisticRegression
# Step 4 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.pipeline import Pipeline
# Step 5 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.preprocessing import OneHotEncoder, StandardScaler

# Step 6 — Instantiate `numeric` with the chosen algorithm/configuration before fitting it to data.
numeric = Pipeline([
    ("impute", SimpleImputer(strategy="median")),
    ("scale", StandardScaler())
])

# Step 7 — Instantiate `categorical` with the chosen algorithm/configuration before fitting it to data.
categorical = Pipeline([
    ("impute", SimpleImputer(strategy="most_frequent")),
    ("encode", OneHotEncoder(handle_unknown="ignore"))
])

# Step 8 — Compute the right-hand expression and store its result in `preprocess` for the next step.
preprocess = ColumnTransformer([
    ("num", numeric, ["age", "income"]),
    ("cat", categorical, ["region"])
])

# Step 9 — Instantiate `model` with the chosen algorithm/configuration before fitting it to data.
model = Pipeline([
    ("preprocess", preprocess),
    ("classifier", LogisticRegression(max_iter=1000))
])

# Step 10 — Display the current value explicitly so the result/state can be inspected during execution.
print(model)
Expected / illustrative result
One estimator now owns numeric imputation/scaling, categorical imputation/encoding, and the classifier.
Interpret the result.

When this object is passed to cross-validation, the numeric medians and scaling statistics and the set of encoded categories are relearned separately from each training fold. The held-out fold is transformed using those training-derived values. That is the key leakage-safety property.

Distinctions & related ideas

Leakage-safe versus leakage-prone workflows

SafeSplit first; fit learned preprocessing and feature selection inside the training fold; transform held-out data with the fitted objects.
LeakyCompute imputation/scaling/feature-selection parameters on the entire dataset before cross-validation.
Data leakageInformation unavailable at genuine prediction time influences model fitting or model-selection decisions.
Target leakageA feature directly or indirectly contains information derived from the outcome or from the future; a pipeline cannot repair this semantic problem automatically.
Use deliberately

Operations that usually belong inside

  • Missing-value imputation.
  • Standardisation, normalisation and learned transforms.
  • Categorical encoding whose categories are inferred from data.
  • Feature selection, PCA and other data-driven dimensionality reduction.
  • The final estimator and tunable preprocessing/model parameters.
Keep outside or decide separately

Design choices before fitting

  • The definition of the prediction target and prediction time.
  • Group, subject or temporal split boundaries.
  • Removal of impossible or post-outcome variables based on domain knowledge.
  • An untouched final test set reserved for final evaluation.
Common mistakes

Where leakage enters

  • Calling fit_transform on all rows and only then performing cross-validation.
  • Selecting “top features” using correlations with the target computed from the full dataset.
  • Applying PCA to all observations before creating folds.
  • Using future measurements, discharge information or post-event fields to predict an earlier outcome.
  • Randomly splitting records from the same person across train and test when the intended generalisation unit is the person.
Verification

Audit the information boundary

  • For every learned preprocessing object, ask: “Which rows were visible when this parameter was estimated?”
  • Run cross-validation on the complete pipeline rather than on preprocessed arrays.
  • Inspect fitted fold behaviour on a small synthetic example where train and validation distributions differ strongly.
  • Keep a final test set untouched until preprocessing, model family and hyperparameters are fixed.
Hands-on practice

Find and remove leakage

Try this:

Take a workflow that performs StandardScaler().fit_transform(X) before cross_val_score. Rewrite it as a Pipeline and compare the validation procedure. Then add a feature-selection step and explain why that step must also remain inside the pipeline.

Pass the raw feature table to cross-validation and make scaling/selection named steps of the estimator. The resampling routine will fit those steps separately in every training fold.
Knowledge check

Identify the safe workflow

Which workflow best protects a cross-validation score from preprocessing leakage?

Quick reference

Remember the logic

Step 1Choose the split strategy first. Decide whether observations may be randomly split or whether time, groups, users, patients or repeated measurements require a constrained splitter.
Step 2Identify all learned transformations. Imputation, scaling, encoding, feature selection and PCA learn parameters from data and therefore belong inside the pipeline.
Step 3Construct column-specific preprocessing. A ColumnTransformer can send numeric and categorical variables through different branches.
Step 4Join preprocessing and estimator. A top-level Pipeline makes the complete transformation-and-model chain a single estimator.
Lesson summary

What to remember

  • A leakage-safe pipeline places every data-dependent preprocessing operation inside the resampling boundary. Imputation values, scaling parameters, category mappings, selected features and dimensionality-reduction components are learned from training data only. The held-out data are transformed with those learned parameters but never participate in fitting them. This is more than code organisation. It protects the validity of model evaluation. If the complete dataset is used to choose features, calculate means or standard deviations, or learn encodings before cross-validation, the validation observations influence the model-building process indirectly. The resulting metric can look better than the performance that would be achieved on genuinely unseen data.
  • Choose the split strategy first. Decide whether observations may be randomly split or whether time, groups, users, patients or repeated measurements require a constrained splitter.
  • Calling fit_transform on all rows and only then performing cross-validation.
  • Keep a final test set untouched until preprocessing, model family and hyperparameters are fixed.