Deeper walkthroughRead Build a Pipeline as a mechanism, not a recipe
Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Define the steps. Each intermediate step must implement a transformation interface such as fit / transform ; the last step is normally an estimator with fit and predict or predict_proba . Stage 2: Call fit(X_train, y_train) . The pipeline fits the first transformer on training data, transforms the data, passes the result to the next transformer, and finally fits the estimator. Stage 3: Call predict(X_new) . The already-fitted transformers modify the new observations using training-derived parameters; they are not refitted on the new data. Final checkpoint: Tune through step names. Parameters are addressed with the step__parameter convention, allowing preprocessing and model settings to be searched together without leaving the leakage-safe workflow.
MechanismFollow the transformation
Define the steps. Each intermediate step must implement a transformation interface such as fit / transform ; the last step is normally an estimator with fit and predict or predict_proba .
Call fit(X_train, y_train) . The pipeline fits the first transformer on training data, transforms the data, passes the result to the next transformer, and finally fits the estimator.
Call predict(X_new) . The already-fitted transformers modify the new observations using training-derived parameters; they are not refitted on the new data.
EvidenceKnow what would convince you
Suppose a binary-classification dataset contains numeric features with a few missing values. We want the median imputation and standardisation to be learned only from training data before fitting logistic regression.
- Inspect pipe.named_steps after fitting and confirm fitted statistics come from training data.
- Compare a cross-validation score from the full pipeline with a correctly implemented manual workflow; they should be consistent.