Data Understanding & Preparation

Preprocessing Playground

Transform messy mixed-type data step by step and see how each choice changes the training representation without letting test information leak backward.

Lab concept guide

What to observe while you experiment

Preprocessing converts raw mixed-type data into the numerical representation consumed by a model. Imputation, scaling and encoding may learn statistics/categories, so they must be fit on training data and then reused unchanged on validation/test data.

MechanismApply each transformation stage separately and inspect the table/feature matrix after it; note which stages learn state during fit.
Failure modeFitting imputation/scaling/encoding on all data before the split or losing track of output feature names and categories.
VerificationCompare train-derived transformation parameters with transformed validation rows and confirm no validation statistic was used to fit them.
Experiment deliberately
Run imputation, scaling and encoding step by step. Predict shape/range changes and identify exactly which learned parameters must be reused at inference.
Leakage-safe by design: imputation, clipping, encoding levels and scaling parameters are learned from the training rows, then applied unchanged to the test rows.
Preparing mixed-type dataset…

Pipeline

Before → after distribution

Transformation diagnostics

Transformed training matrix preview