Experimental Design & Validation · Lesson 50

Train Validation Test

Train Validation Test is a model-validation design.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What Train Validation Test actually means

Train Validation Test is a model-validation design. Validation estimates how a trained procedure will generalise to new cases, so the split must reproduce the independence structure and time/entity boundaries expected at deployment.

Train Validation Test matters because validation is the evidence for generalisation. The split strategy must reproduce the independence, grouping and temporal structure of future use; otherwise a high score can be an artefact of leakage.

Deeper walkthrough

Read Train Validation Test as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Define the prediction unit and what “new” means: new row, subject, group or future time. Stage 2: Reserve test data for final evaluation when feasible. Stage 3: Use cross-validation inside development to compare models/hyperparameters. Final checkpoint: Use group- or time-aware splits when ordinary random splitting would leak related/future information.

Mechanism

Follow the transformation

Define the prediction unit and what “new” means: new row, subject, group or future time.

Reserve test data for final evaluation when feasible.

Use cross-validation inside development to compare models/hyperparameters.

Evidence

Know what would convince you

  • Write down the prediction unit and which records are allowed to coexist across train/validation/test before splitting.
  • Inspect fold/group/time indices directly and assert that forbidden overlap is zero.
Useful distinctionHoldout: One train/test split; simple but higher variance.
Visual demonstration of Train Validation Test
Visual demonstration: use the diagram to trace the main objects and state changes involved in Train Validation Test.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Define the prediction unit and what “new” means

Define the prediction unit and what “new” means: new row, subject, group or future time. Treat the output from Train Validation Test as evidence to inspect: confirm its type, shape, range or units and connect it back to the input that produced it.

Output focus: inspect both the value and its shape/type/meaning before treating it as a trustworthy result.
How it works

Trace the mechanism step by step

  1. Define the prediction unit and what “new” means: new row, subject, group or future time.
  2. Reserve test data for final evaluation when feasible.
  3. Use cross-validation inside development to compare models/hyperparameters.
  4. Keep all preprocessing, feature selection and tuning inside each training fold.
  5. Use group- or time-aware splits when ordinary random splitting would leak related/future information.
Worked demonstration

Make the concept concrete

Demonstration

Python / scikit-learn example

# Step 1 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.model_selection import KFold
# Step 2 — Compute the right-hand expression and store its result in `X` for the next step.
X = list(range(6))
# Step 3 — Iterate through the collection so the indented block is applied once for each item.
for fold,(tr,va) in enumerate(KFold(3,shuffle=True,random_state=1).split(X),1):
    # Step 4 — Display the current value explicitly so the result/state can be inspected during execution.
    print(fold, "train", tr.tolist(), "valid", va.tolist())
Expected / illustrative result
Each sample is validation exactly once across three folds; metrics are aggregated across folds.
Interpret the result.

For Train Validation Test, connect the reported result to the exact training/validation/prediction step that produced it and check one prediction, fold or metric component independently.

Distinctions & related ideas

Know what this is — and what it is not

HoldoutOne train/test split; simple but higher variance.
K-foldEach fold serves once as validation; assumes rows are exchangeable.
StratifiedPreserves class proportions approximately.
Group K-foldKeeps all rows from the same group together.
Time-series splitTrains on past and validates on later data.
Nested CVOuter loop estimates generalisation; inner loop selects/tunes.
Use deliberately

When it is appropriate

Use Train Validation Test when it reflects how genuinely unseen cases will arrive and keeps every learned choice inside the training portion of each evaluation split.

Boundary conditions

When to stop or reconsider

Choose a different split strategy when observations share subjects/groups, have temporal order, or otherwise violate independent random splitting assumptions.

Common mistakes

Failure modes to recognise

  • Letting the final test set influence feature engineering, model choice, tuning or threshold selection.
  • Splitting related groups or future/past records in a way that leaks information across folds.
  • Reporting one lucky split without examining variability or preserving the exact split logic.
Verification

How to check the result

  • Write down the prediction unit and which records are allowed to coexist across train/validation/test before splitting.
  • Inspect fold/group/time indices directly and assert that forbidden overlap is zero.
  • Run preprocessing/tuning inside each training fold and reserve the final test set for one final evaluation.
Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of Train Validation Test. First define the prediction unit and what “new” means: new row, subject, group or future time. Then reserve test data for final evaluation when feasible. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Draw the rows/groups/times as blocks and label exactly which block trains, validates and tests each step. Check for any information path crossing the boundary.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from Train Validation Test, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Define the prediction unit and what “new” means: new row, subject, group or future time.
Step 2Reserve test data for final evaluation when feasible.
Step 3Use cross-validation inside development to compare models/hyperparameters.
Step 4Keep all preprocessing, feature selection and tuning inside each training fold.
Lesson summary

What to remember

  • Train Validation Test is a model-validation design. Validation estimates how a trained procedure will generalise to new cases, so the split must reproduce the independence structure and time/entity boundaries expected at deployment.
  • Define the prediction unit and what “new” means: new row, subject, group or future time.
  • Letting the final test set influence feature engineering, model choice, tuning or threshold selection.
  • Write down the prediction unit and which records are allowed to coexist across train/validation/test before splitting.