Preprocessing · Lesson 41

Pipelines

A modelling pipeline is an ordered, fitted procedure that transforms raw features and then produces predictions.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What Pipelines actually means

A modelling pipeline is an ordered, fitted procedure that transforms raw features and then produces predictions. In data science, the pipeline is the unit that should be cross-validated, persisted and reused at inference time.

If preprocessing is performed outside the pipeline, it is easy to leak validation information, forget a transformation at deployment, or apply inconsistent column logic to new data.

Deeper walkthrough

Read Pipelines as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Define every learned preprocessing step. Stage 2: Chain steps in execution order. Stage 3: Fit the whole pipeline on training data. Final checkpoint: Persist and deploy the fitted pipeline rather than the estimator alone.

Mechanism

Follow the transformation

Define every learned preprocessing step.

Chain steps in execution order.

Fit the whole pipeline on training data.

Evidence

Know what would convince you

  • Fit the transformation on training data only and confirm validation/test rows are transformed without refitting.
  • Inspect transformed shape, feature names/ranges and several rows against the raw inputs.
Useful distinctionStandalone estimator: Assumes features are already in the correct representation.
Visual demonstration of Pipelines
Visual demonstration: use the diagram to trace the main objects and state changes involved in Pipelines.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Define every learned preprocessing step

Define every learned preprocessing step. This is an input-preparation stage for Pipelines. Verify the relevant type, shape, units, keys, missingness or assumptions before later steps depend on them.

Input focus: confirm the data/object, units, type, shape and assumptions before the next operation depends on them.
How it works

Trace the mechanism step by step

  1. Define every learned preprocessing step.
  2. Chain steps in execution order.
  3. Fit the whole pipeline on training data.
  4. Evaluate the whole pipeline on held-out data.
  5. Persist and deploy the fitted pipeline rather than the estimator alone.
Worked demonstration

Make the concept concrete

Demonstration

Python / scikit-learn example

# Step 1 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.pipeline import Pipeline
# Step 2 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.impute import SimpleImputer
# Step 3 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.preprocessing import StandardScaler
# Step 4 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.linear_model import LogisticRegression
# Step 5 — Create `pipe` as the scaling object; its parameters will be learned from training data.
pipe = Pipeline([("impute", SimpleImputer(strategy="median")), ("scale", StandardScaler()), ("model", LogisticRegression())])
# Step 6 — Display the current value explicitly so the result/state can be inspected during execution.
print(pipe)
Expected / illustrative result
The same imputation and scaling learned on training data are automatically applied before every prediction.
Interpret the result.

For Pipelines, trace representative source rows/columns into the result and reconcile row counts, dtypes, keys or missing values that the operation could change.

Distinctions & related ideas

Know what this is — and what it is not

Standalone estimatorAssumes features are already in the correct representation.
PipelineOwns preprocessing + estimator as one reproducible procedure.
Manual notebook sequenceCan hide state and make inference inconsistent.
Use deliberately

When it is appropriate

Use Pipelines when the model/analysis requires a deliberate representation of raw features and the transformation can be fit without leaking future or held-out information.

Boundary conditions

When to stop or reconsider

Avoid transformations that are unnecessary for the chosen model, cannot be reproduced at inference time, or learn from data that should remain held out.

Common mistakes

Failure modes to recognise

  • Fitting scaling, encoding, imputation or feature selection on the full dataset before validation.
  • Creating feature values whose meaning changes between training and inference.
  • Ignoring output shape/feature names and losing track of what the transformed columns represent.
Verification

How to check the result

  • Fit the transformation on training data only and confirm validation/test rows are transformed without refitting.
  • Inspect transformed shape, feature names/ranges and several rows against the raw inputs.
  • Place the transformation in a pipeline and cross-validate the complete workflow, not a preprocessed snapshot.
Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of Pipelines. First define every learned preprocessing step. Then chain steps in execution order. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Use a tiny train/validation split. Fit only on train, print the transformed shape/values, and confirm validation transformation does not update learned state.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from Pipelines, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Define every learned preprocessing step.
Step 2Chain steps in execution order.
Step 3Fit the whole pipeline on training data.
Step 4Evaluate the whole pipeline on held-out data.
Lesson summary

What to remember

  • Pipelines belongs to Python code organisation and dependency management. Modules split source into importable files, packages group modules, virtual environments isolate dependencies, and requirement metadata makes the software environment reproducible.
  • Define every learned preprocessing step.
  • Fitting scaling, encoding, imputation or feature selection on the full dataset before validation.
  • Fit the transformation on training data only and confirm validation/test rows are transformed without refitting.