Training, Validation & Model Selection

ML Pipeline Builder

Connect data preparation, feature work, modelling and evaluation into one inspectable workflow.

Lab concept guide

What to observe while you experiment

A machine-learning pipeline is an ordered data contract from raw features through preprocessing/feature work to a fitted estimator and evaluation. Putting learned transformations inside the pipeline keeps training and validation boundaries intact.

MechanismConnect each stage in order and identify which stages learn state during fit versus only transform/predict during inference.
Failure modePreprocessing the full dataset before splitting/cross-validation or creating an inference pipeline that omits training-time transformations.
VerificationFit the complete pipeline on training data, run raw validation rows through it without refitting, and inspect shape/feature consistency at stage boundaries.
Experiment deliberately
Build a minimal impute → scale/encode → model pipeline. Predict which stages learn parameters, then cross-validate the entire pipeline.
Correct order matters. Every preprocessing step is fitted on training data only, then reused on validation/test data. The test set is never used to choose the pipeline.
Ready.

Pipeline

Dataset → clean → transform → engineer → select → model → evaluate.

Split Geometry with Model

The observations stay fixed while the selected classifier separator is revealed over the train/validation/test split.

Evidence by split

Use validation evidence to choose; use test evidence only for the final check.