6 · Splitting, Validation & Experiment Design

Data Leakage & Experimental Integrity

Data leakage occurs when information unavailable at real prediction time influences training, feature construction, selection, tuning or evaluation. The topic is split into focused lessons so definitions, implementation decisions, diagnostics and common failure modes can be learned separately.

How to use this topic

Learn the mechanism one decision at a time

Work through the lessons in order if the topic is new. If you already know the basics, open the specific leaf lesson that matches the operation, diagnostic or failure mode you need.

1Definition→
2Mechanism→
3Example→
4Diagnostic→
5Decision
01
Preprocessing leakageScaling, imputation, PCA or feature selection fitted on the full dataset exposes validation/test distribution information. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
02
Target leakageFeatures directly or indirectly encode outcomes that would not be known at prediction time. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
03
Group leakageRecords from the same patient, customer, device, recording or source appear in both train and test sets. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
04
Temporal leakageFuture observations influence historical features, normalization, imputation or model tuning. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
05
Duplicate and pipeline leakageNear-duplicates, augmented copies or cached derived features can cross split boundaries unless lineage is tracked. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.