Preprocessing · Lesson 38

Categorical Encoding

Categorical Encoding is a preprocessing operation that changes feature representation before modelling.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What Categorical Encoding actually means

Categorical Encoding is a preprocessing operation that changes feature representation before modelling. Preprocessing is part of the learned pipeline because statistics such as means, categories or quantiles must be estimated on training data and then applied unchanged to validation/test data.

Categorical Encoding matters because preprocessing defines the representation a model actually sees. Learned transformations must be fitted inside the training boundary so their parameters do not leak information from validation or test cases.

Deeper walkthrough

Read Categorical Encoding as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Split the data before fitting learned preprocessing. Stage 2: Apply transformations separately by column type. Stage 3: Fit preprocessing only on the training fold. Final checkpoint: Bundle preprocessing and estimator in one pipeline so cross-validation repeats the correct sequence automatically.

Mechanism

Follow the transformation

Split the data before fitting learned preprocessing.

Apply transformations separately by column type.

Fit preprocessing only on the training fold.

Evidence

Know what would convince you

  • Fit the transformation on training data only and confirm validation/test rows are transformed without refitting.
  • Inspect transformed shape, feature names/ranges and several rows against the raw inputs.
Useful distinctionStandardisation: (x - training mean) / training standard deviation.
Visual demonstration of Categorical Encoding
Visual demonstration: use the diagram to trace the main objects and state changes involved in Categorical Encoding.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Split the data before fitting learned…

Split the data before fitting learned preprocessing. At this stage of Categorical Encoding, keep the incoming data or object separate from the learned parameter, transformed object, or statistic so the change can be reproduced and independently checked.

Transformation focus: keep the input and produced parameters/result separate so the change is observable and reproducible.
How it works

Trace the mechanism step by step

  1. Split the data before fitting learned preprocessing.
  2. Apply transformations separately by column type.
  3. Fit preprocessing only on the training fold.
  4. Transform validation/test with the stored training parameters.
  5. Bundle preprocessing and estimator in one pipeline so cross-validation repeats the correct sequence automatically.
Worked demonstration

Make the concept concrete

Demonstration

Python / scikit-learn example

# Step 1 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.compose import ColumnTransformer
# Step 2 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.preprocessing import OneHotEncoder, StandardScaler
# Step 3 — Compute the right-hand expression and store its result in `pre` for the next step.
pre = ColumnTransformer([
    ("num", StandardScaler(), ["age","income"]),
    ("cat", OneHotEncoder(handle_unknown="ignore"), ["region"])
])
# Step 4 — Display the current value explicitly so the result/state can be inspected during execution.
print(pre)
Expected / illustrative result
Numeric and categorical columns receive different transformations in one fitted preprocessing object.
Interpret the result.

For Categorical Encoding, trace representative source rows/columns into the result and reconcile row counts, dtypes, keys or missing values that the operation could change.

Distinctions & related ideas

Know what this is — and what it is not

Standardisation(x - training mean) / training standard deviation.
Robust scalingUses median and quantile range; less influenced by extremes.
One-hot encodingCreates indicator columns for categories without imposing numeric order.
PipelineChains preprocessing and estimator so fitting/evaluation stay leakage-safe.
Use deliberately

When it is appropriate

Use Categorical Encoding when the model/analysis requires a deliberate representation of raw features and the transformation can be fit without leaking future or held-out information.

Boundary conditions

When to stop or reconsider

Avoid transformations that are unnecessary for the chosen model, cannot be reproduced at inference time, or learn from data that should remain held out.

Common mistakes

Failure modes to recognise

  • Fitting scaling, encoding, imputation or feature selection on the full dataset before validation.
  • Creating feature values whose meaning changes between training and inference.
  • Ignoring output shape/feature names and losing track of what the transformed columns represent.
Verification

How to check the result

  • Fit the transformation on training data only and confirm validation/test rows are transformed without refitting.
  • Inspect transformed shape, feature names/ranges and several rows against the raw inputs.
  • Place the transformation in a pipeline and cross-validate the complete workflow, not a preprocessed snapshot.
Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of Categorical Encoding. First split the data before fitting learned preprocessing. Then apply transformations separately by column type. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Use a tiny train/validation split. Fit only on train, print the transformed shape/values, and confirm validation transformation does not update learned state.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from Categorical Encoding, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Split the data before fitting learned preprocessing.
Step 2Apply transformations separately by column type.
Step 3Fit preprocessing only on the training fold.
Step 4Transform validation/test with the stored training parameters.
Lesson summary

What to remember

  • Categorical Encoding is a preprocessing operation that changes feature representation before modelling. Preprocessing is part of the learned pipeline because statistics such as means, categories or quantiles must be estimated on training data and then applied unchanged to validation/test data.
  • Split the data before fitting learned preprocessing.
  • Fitting scaling, encoding, imputation or feature selection on the full dataset before validation.
  • Fit the transformation on training data only and confirm validation/test rows are transformed without refitting.