Experiment Design · Flagship experience

Train, Validation & Test

Why do we need data the model has never seen?

Start here

Why do we need data the model has never seen?

Training measures fit; validation guides choices; testing estimates final generalisation. Mixing these roles lets decisions adapt to the answer key.

Building interactive view…
Understand

Build the mental model

Training measures fit; validation guides choices; testing estimates final generalisation. Mixing these roles lets decisions adapt to the answer key. Split by the deployment unit—person, customer, device, time—not blindly by row.

Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Train

This is a learning/estimation stage. Separate the data supplied to the algorithm from the parameters or structure it learns, and keep validation information outside the fit. Technical context for Train, Validation & Test: Repeated model selection on a validation set gradually overfits that validation evidence. The final test set should remain untouched until choices are frozen.

Practitioner checkpoint: Split by the deployment unit—person, customer, device, time—not blindly by row.
What happens if…?

Break the assumption deliberately

Tune repeatedly on the test set and watch the “final” estimate stop being independent.

Move the control and explain what you expect before reading the visual.

Technical lens

Formalise what the visual is doing

Repeated model selection on a validation set gradually overfits that validation evidence. The final test set should remain untouched until choices are frozen.

Technical questionUse a tiny case to make the mechanism observable. Repeated model selection on a validation set gradually overfits that validation evidence. The final test set should remain untouched until choices are frozen. Verify one intermediate quantity, state change or mapping independently; then predict the consequence of this change: Tune repeatedly on the test set and watch the “final” estimate stop being independent.
Practitioner lens

Use it responsibly

Split by the deployment unit—person, customer, device, time—not blindly by row.

Transfer testTransfer this idea to a new example and justify each decision using this practitioner rule: Split by the deployment unit—person, customer, device, time—not blindly by row. Then explain what should change if you deliberately test: Tune repeatedly on the test set and watch the “final” estimate stop being independent.
Worked exploration

Use the visual as an experiment, not decoration

Split 100 labelled cases into training, validation and test sets. Fit on training, choose a threshold on validation, and evaluate once on test. Then imagine tuning repeatedly on test and explain why it stops being an unbiased final check.

Technical lens

Repeated model selection on a validation set gradually overfits that validation evidence. The final test set should remain untouched until choices are frozen.

Practitioner check

Split by the deployment unit—person, customer, device, time—not blindly by row.

Prediction before interaction
Tune repeatedly on the test set and watch the “final” estimate stop being independent.
Exploration walkthrough

Turn the interaction into an evidence trail

Split 100 labelled cases into training, validation and test sets. Fit on training, choose a threshold on validation, and evaluate once on test. Then imagine tuning repeatedly on test and explain why it stops being an unbiased final check. Before moving the control, state your prediction. After the visual changes, name the specific state, statistic, boundary or mapping that changed and explain why that change is consistent—or inconsistent—with your prediction.

  • Record one observable quantity before the interaction and the same quantity afterwards.
  • Change one factor at a time so the causal effect of the control is inspectable.
  • Use an edge or failure case to discover where the concept stops behaving as the simple story suggests.
Reference depth

Open the complete material

The flagship experience is the map. These pages contain the roads.

Deep Learning Hub lessons

Three-way train-validation-testTrain / Validation / Test Design

Related models & simulations

Use the Concept Atlas for related methods and adjacent concepts.

Continue this exact concept

Choose depth, practice or application.

These destinations are explicitly mapped to Train, Validation & Test; they are not generic landing-page fallbacks.