Capstone Project · Lesson 90

Build a Pipeline

A machine-learning pipeline is an ordered chain of transformations and a final estimator that behaves like one model.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What a machine-learning pipeline is

A machine-learning pipeline is an ordered chain of transformations and a final estimator that behaves like one model. Instead of manually imputing missing values, scaling features, fitting a classifier and remembering to repeat the same operations on new data, a pipeline records those steps as a single executable object.

This matters most during validation. Each preprocessing step must be learned from the training portion of a fold and then applied to that fold's validation portion. When preprocessing is fitted before cross-validation, information from the validation data can influence the transformation and produce overly optimistic scores. A pipeline keeps the fit/transform boundary inside each training fold.

Deeper walkthrough

Read Build a Pipeline as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Define the steps. Each intermediate step must implement a transformation interface such as fit / transform ; the last step is normally an estimator with fit and predict or predict_proba . Stage 2: Call fit(X_train, y_train) . The pipeline fits the first transformer on training data, transforms the data, passes the result to the next transformer, and finally fits the estimator. Stage 3: Call predict(X_new) . The already-fitted transformers modify the new observations using training-derived parameters; they are not refitted on the new data. Final checkpoint: Tune through step names. Parameters are addressed with the step__parameter convention, allowing preprocessing and model settings to be searched together without leaving the leakage-safe workflow.

Mechanism

Follow the transformation

Define the steps. Each intermediate step must implement a transformation interface such as fit / transform ; the last step is normally an estimator with fit and predict or predict_proba .

Call fit(X_train, y_train) . The pipeline fits the first transformer on training data, transforms the data, passes the result to the next transformer, and finally fits the estimator.

Call predict(X_new) . The already-fitted transformers modify the new observations using training-derived parameters; they are not refitted on the new data.

Evidence

Know what would convince you

Suppose a binary-classification dataset contains numeric features with a few missing values. We want the median imputation and standardisation to be learned only from training data before fitting logistic regression.

  • Inspect pipe.named_steps after fitting and confirm fitted statistics come from training data.
  • Compare a cross-validation score from the full pipeline with a correctly implemented manual workflow; they should be consistent.
Useful distinctionManual preprocessing: Can be correct, but the programmer must enforce the train/test boundary for every transformation and every resampling fold.
Visual demonstration of Build a Pipeline
Visual demonstration: use the diagram to trace the main objects and state changes involved in Build a Pipeline.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Define the steps

Define the steps. Each intermediate step must implement a transformation interface such as fit / transform ; the last step is normally an estimator with fit and predict or predict_proba .

Transformation focus: keep the input and produced parameters/result separate so the change is observable and reproducible.
How it works

Fit and predict follow different paths

  1. Define the steps. Each intermediate step must implement a transformation interface such as fit/transform; the last step is normally an estimator with fit and predict or predict_proba.
  2. Call fit(X_train, y_train). The pipeline fits the first transformer on training data, transforms the data, passes the result to the next transformer, and finally fits the estimator.
  3. Call predict(X_new). The already-fitted transformers modify the new observations using training-derived parameters; they are not refitted on the new data.
  4. Use it inside cross-validation. The entire pipeline is cloned and fitted separately in each fold, so medians, means, scaling parameters and learned model coefficients come only from that fold's training data.
  5. Tune through step names. Parameters are addressed with the step__parameter convention, allowing preprocessing and model settings to be searched together without leaving the leakage-safe workflow.
Worked demonstration

Imputation + scaling + logistic regression

Suppose a binary-classification dataset contains numeric features with a few missing values. We want the median imputation and standardisation to be learned only from training data before fitting logistic regression.

Demonstration

Build one estimator from three steps

# Step 1 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.impute import SimpleImputer
# Step 2 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.linear_model import LogisticRegression
# Step 3 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.pipeline import Pipeline
# Step 4 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.preprocessing import StandardScaler

# Step 5 — Instantiate `pipe` with the chosen algorithm/configuration before fitting it to data.
pipe = Pipeline([
    ("impute", SimpleImputer(strategy="median")),
    ("scale", StandardScaler()),
    ("model", LogisticRegression(max_iter=1000))
])

# Step 6 — Display the current value explicitly so the result/state can be inspected during execution.
print(pipe)
Expected / illustrative result
Pipeline(steps=[('impute', SimpleImputer(strategy='median')),
                ('scale', StandardScaler()),
                ('model', LogisticRegression(max_iter=1000))])
Interpret the result.

The object has not learned anything yet; it is a recipe. After pipe.fit(X_train, y_train), the imputer stores training medians, the scaler stores training means and standard deviations, and logistic regression stores learned coefficients. pipe.predict(X_test) applies those stored transformations to the test data before prediction.

Validation consequence: placing this pipeline inside cross-validation makes each fold learn its own imputation and scaling parameters rather than leaking information from the held-out fold.
Distinctions & related ideas

Pipeline versus manual preprocessing

Manual preprocessingCan be correct, but the programmer must enforce the train/test boundary for every transformation and every resampling fold.
PipelineEncapsulates preprocessing and the estimator so fitting, validation and prediction follow the same declared sequence.
ColumnTransformerRuns different preprocessing branches for different columns, such as scaling numeric variables and encoding categorical variables; it can itself be a pipeline step.
Grid/Random searchCan tune nested parameters with names such as model__C while refitting the whole pipeline inside each validation split.
Use deliberately

When a pipeline is especially useful

  • Preprocessing has learned parameters: imputation, scaling, feature selection, dimensionality reduction or encoding.
  • You use cross-validation or hyperparameter optimisation.
  • The same preprocessing must be reproduced at inference time.
  • You want a deployable object that captures preprocessing and prediction together.
Boundary conditions

What a pipeline does not solve automatically

  • It cannot decide whether the train/test split itself is appropriate for time, groups or repeated subjects.
  • It cannot prevent leakage from features that already contain future or target-derived information.
  • Custom transformers can still leak if their implementation reads data outside the supplied training fold.
Common mistakes

Failure modes to recognise

  • Scaling or imputing the complete dataset before the train/test split.
  • Calling fit_transform separately on the test set, which learns a different transformation from test data.
  • Performing feature selection before cross-validation instead of placing selection inside the pipeline.
  • Using ordinary random folds when time, group or subject structure requires a specialised splitter.
Verification

How to check the result

  • Inspect pipe.named_steps after fitting and confirm fitted statistics come from training data.
  • Compare a cross-validation score from the full pipeline with a correctly implemented manual workflow; they should be consistent.
  • Run prediction on raw held-out rows—not manually preprocessed copies—to confirm the pipeline performs the transformations itself.
Hands-on practice

Build and inspect a pipeline

Try this:

Create a pipeline with median imputation, standardisation and logistic regression. Fit it on a training split, inspect the imputer's statistics_ and the scaler's mean_, then verify that calling predict on raw test rows succeeds without separate preprocessing.

Use pipe.named_steps["impute"].statistics_ and pipe.named_steps["scale"].mean_ after fit. The test data should be passed directly to pipe.predict.
Knowledge check

Check the validation boundary

Why should scaling normally be placed inside a pipeline used by cross-validation?

Quick reference

Remember the logic

Step 1Define the steps. Each intermediate step must implement a transformation interface such as fit / transform ; the last step is normally an estimator with fit and predict or predict_proba .
Step 2Call fit(X_train, y_train) . The pipeline fits the first transformer on training data, transforms the data, passes the result to the next transformer, and finally fits the estimator.
Step 3Call predict(X_new) . The already-fitted transformers modify the new observations using training-derived parameters; they are not refitted on the new data.
Step 4Use it inside cross-validation. The entire pipeline is cloned and fitted separately in each fold, so medians, means, scaling parameters and learned model coefficients come only from that fold's training data.
Lesson summary

What to remember

  • A machine-learning pipeline is an ordered chain of transformations and a final estimator that behaves like one model. Instead of manually imputing missing values, scaling features, fitting a classifier and remembering to repeat the same operations on new data, a pipeline records those steps as a single executable object. This matters most during validation. Each preprocessing step must be learned from the training portion of a fold and then applied to that fold's validation portion. When preprocessing is fitted before cross-validation, information from the validation data can influence the transformation and produce overly optimistic scores. A pipeline keeps the fit/transform boundary inside each training fold.
  • Define the steps. Each intermediate step must implement a transformation interface such as fit / transform ; the last step is normally an estimator with fit and predict or predict_proba .
  • Scaling or imputing the complete dataset before the train/test split.
  • Run prediction on raw held-out rows—not manually preprocessed copies—to confirm the pipeline performs the transformations itself.