Cross-Validation & Data Splitting
Validation estimates how a model will behave on unseen data. The split strategy must mirror the real independence structure of samples, subjects, groups and time. The topic is split into focused lessons so definitions, implementation decisions, diagnostics and common failure modes can be learned separately.
Learn the mechanism one decision at a time
Work through the lessons in order if the topic is new. If you already know the basics, open the specific leaf lesson that matches the operation, diagnostic or failure mode you need.
1Definition→
2Mechanism→
3Example→
4Diagnostic→
5Decision
Holdout and K-FoldHoldout reserves one test partition. K-Fold rotates the validation partition through K disjoint folds so every observation is validated once. Repeated K-Fold reduces dependence on one fold assignment. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
02StratificationStratified splitting preserves class proportions. It is valuable for classification, especially when minority classes are small, but it does not solve subject or group dependence. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
03Grouped validationGroup K-Fold keeps all records from the same subject, patient, customer, device or site together. Stratified Group K-Fold additionally tries to preserve class balance without breaking groups. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
04Time-aware validationTime-Series Split, rolling windows and expanding windows preserve chronological order. Future observations must never influence training features, preprocessing or tuning for earlier predictions. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
05Nested cross-validationNested CV separates hyperparameter selection from performance estimation: an inner loop tunes the model and an outer loop estimates generalisation. It is computationally expensive but reduces optimisation bias. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.