Start hereWhy do we need data the model has never seen?
Training measures fit; validation guides choices; testing estimates final generalisation. Mixing these roles lets decisions adapt to the answer key.
Building interactive view…
Technical lensFormalise what the visual is doing
Repeated model selection on a validation set gradually overfits that validation evidence. The final test set should remain untouched until choices are frozen.
Technical questionUse a tiny case to make the mechanism observable. Repeated model selection on a validation set gradually overfits that validation evidence. The final test set should remain untouched until choices are frozen. Verify one intermediate quantity, state change or mapping independently; then predict the consequence of this change: Tune repeatedly on the test set and watch the “final” estimate stop being independent.
Practitioner lensUse it responsibly
Split by the deployment unit—person, customer, device, time—not blindly by row.
Transfer testTransfer this idea to a new example and justify each decision using this practitioner rule: Split by the deployment unit—person, customer, device, time—not blindly by row. Then explain what should change if you deliberately test: Tune repeatedly on the test set and watch the “final” estimate stop being independent.
Worked explorationUse the visual as an experiment, not decoration
Split 100 labelled cases into training, validation and test sets. Fit on training, choose a threshold on validation, and evaluate once on test. Then imagine tuning repeatedly on test and explain why it stops being an unbiased final check.
Technical lens
Repeated model selection on a validation set gradually overfits that validation evidence. The final test set should remain untouched until choices are frozen.
Practitioner check
Split by the deployment unit—person, customer, device, time—not blindly by row.
Prediction before interactionTune repeatedly on the test set and watch the “final” estimate stop being independent.
Exploration walkthroughTurn the interaction into an evidence trail
Split 100 labelled cases into training, validation and test sets. Fit on training, choose a threshold on validation, and evaluate once on test. Then imagine tuning repeatedly on test and explain why it stops being an unbiased final check. Before moving the control, state your prediction. After the visual changes, name the specific state, statistic, boundary or mapping that changed and explain why that change is consistent—or inconsistent—with your prediction.
- Record one observable quantity before the interaction and the same quantity afterwards.
- Change one factor at a time so the causal effect of the control is inspectable.
- Use an edge or failure case to discover where the concept stops behaving as the simple story suggests.