Preprocessing Playground
Transform messy mixed-type data step by step and see how each choice changes the training representation without letting test information leak backward.
What to observe while you experiment
Preprocessing converts raw mixed-type data into the numerical representation consumed by a model. Imputation, scaling and encoding may learn statistics/categories, so they must be fit on training data and then reused unchanged on validation/test data.
Experiment deliberately
Run imputation, scaling and encoding step by step. Predict shape/range changes and identify exactly which learned parameters must be reused at inference.
Leakage-safe by design: imputation, clipping, encoding levels and scaling parameters are learned from the training rows, then applied unchanged to the test rows.
Preparing mixed-type dataset…