Data Preparation for ML · Lesson 20

Pipelines and Columntransformer

A scikit-learn Pipeline chains sequential transformations with an estimator, while ColumnTransformer applies different preprocessing to different feature subsets before combining them.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What Pipelines and Columntransformer actually means

A scikit-learn Pipeline chains sequential transformations with an estimator, while ColumnTransformer applies different preprocessing to different feature subsets before combining them. Together they make heterogeneous preprocessing reproducible and leakage-safe.

Real datasets mix numeric, categorical and other columns. Keeping their transformations inside one fitted object ensures training, validation and inference use the same logic and learned parameters.

Deeper walkthrough

Read Pipelines and Columntransformer as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Partition columns by semantic type or preprocessing need. Stage 2: Define numeric and categorical transformers separately. Stage 3: Combine them with ColumnTransformer. Final checkpoint: Fit the full pipeline only on training data and reuse it unchanged for inference.

Mechanism

Follow the transformation

Partition columns by semantic type or preprocessing need.

Define numeric and categorical transformers separately.

Combine them with ColumnTransformer.

Evidence

Know what would convince you

  • Fit the transformation on training data only and confirm validation/test rows are transformed without refitting.
  • Inspect transformed shape, feature names/ranges and several rows against the raw inputs.
Useful distinctionPipeline: Sequential steps applied to the same evolving feature matrix.
Visual demonstration of Pipelines and Columntransformer
Visual demonstration: use the diagram to trace the main objects and state changes involved in Pipelines and Columntransformer.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Partition columns by semantic type or…

Partition columns by semantic type or preprocessing need. For Pipelines and Columntransformer, identify the exact state before this stage, the operation or rule applied here, and the observable state afterwards so the mechanism remains inspectable.

State focus: identify exactly what changed at this stage and what observable evidence confirms that change.
How it works

Trace the mechanism step by step

  1. Partition columns by semantic type or preprocessing need.
  2. Define numeric and categorical transformers separately.
  3. Combine them with ColumnTransformer.
  4. Attach the estimator as the final Pipeline step.
  5. Fit the full pipeline only on training data and reuse it unchanged for inference.
Worked demonstration

Make the concept concrete

Demonstration

Python / scikit-learn example

# Step 1 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.compose import ColumnTransformer
# Step 2 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.pipeline import Pipeline
# Step 3 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.preprocessing import OneHotEncoder, StandardScaler
# Step 4 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.linear_model import LogisticRegression
# Step 5 — Create `pre` as the scaling object; its parameters will be learned from training data.
pre = ColumnTransformer([("num", StandardScaler(), ["age","income"]), ("cat", OneHotEncoder(handle_unknown="ignore"), ["region"])])
# Step 6 — Instantiate `pipe` with the chosen algorithm/configuration before fitting it to data.
pipe = Pipeline([("pre", pre), ("model", LogisticRegression())])
# Step 7 — Display the current value explicitly so the result/state can be inspected during execution.
print(pipe)
Expected / illustrative result
Numeric and categorical columns are transformed appropriately, then passed into one estimator through a single fitted pipeline.
Interpret the result.

For Pipelines and Columntransformer, trace representative source rows/columns into the result and reconcile row counts, dtypes, keys or missing values that the operation could change.

Distinctions & related ideas

Know what this is — and what it is not

PipelineSequential steps applied to the same evolving feature matrix.
ColumnTransformerParallel transformations applied to selected columns, then concatenated.
Manual preprocessingEasy to make train/test logic inconsistent or leaky.
Use deliberately

When it is appropriate

Use Pipelines and Columntransformer when the model/analysis requires a deliberate representation of raw features and the transformation can be fit without leaking future or held-out information.

Boundary conditions

When to stop or reconsider

Avoid transformations that are unnecessary for the chosen model, cannot be reproduced at inference time, or learn from data that should remain held out.

Common mistakes

Failure modes to recognise

  • Fitting scaling, encoding, imputation or feature selection on the full dataset before validation.
  • Creating feature values whose meaning changes between training and inference.
  • Ignoring output shape/feature names and losing track of what the transformed columns represent.
Verification

How to check the result

  • Fit the transformation on training data only and confirm validation/test rows are transformed without refitting.
  • Inspect transformed shape, feature names/ranges and several rows against the raw inputs.
  • Place the transformation in a pipeline and cross-validate the complete workflow, not a preprocessed snapshot.
Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of Pipelines and Columntransformer. First partition columns by semantic type or preprocessing need. Then define numeric and categorical transformers separately. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Use a tiny train/validation split. Fit only on train, print the transformed shape/values, and confirm validation transformation does not update learned state.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from Pipelines and Columntransformer, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Partition columns by semantic type or preprocessing need.
Step 2Define numeric and categorical transformers separately.
Step 3Combine them with ColumnTransformer.
Step 4Attach the estimator as the final Pipeline step.
Lesson summary

What to remember

  • Pipelines and Columntransformer belongs to Python code organisation and dependency management. Modules split source into importable files, packages group modules, virtual environments isolate dependencies, and requirement metadata makes the software environment reproducible.
  • Partition columns by semantic type or preprocessing need.
  • Fitting scaling, encoding, imputation or feature selection on the full dataset before validation.
  • Fit the transformation on training data only and confirm validation/test rows are transformed without refitting.