Training, Validation & Model Selection

Overfitting / Underfitting Lab

Manipulate data size, noise and model complexity to see underfitting, useful fit and overfitting emerge.

Lab concept guide

What to observe while you experiment

Generalisation depends on the match between model capacity, data amount/noise and the underlying signal. Underfitting shows high error even on training data; overfitting shows a widening train–validation gap as the model captures idiosyncrasies.

MechanismChange model complexity, sample size or noise one at a time and compare training with validation behaviour.
Failure modeCalling any high training score “good fit” without checking unseen data, or diagnosing overfitting from a single noisy split.
VerificationTrack train and validation metrics across complexity and repeat/inspect variability; identify the region where added complexity stops improving validation.
Experiment deliberately
Increase complexity from too simple to very flexible. Predict the training curve and validation curve shape before running the sweep.
Bias–variance is visible. Increase polynomial degree or reduce data to see training error fall while validation error can rise.
Ready.
UnderfittingBoth training and validation error are high.
Useful fitValidation error is near its minimum and the gap is controlled.
OverfittingTraining error keeps falling while validation error rises.

Current model fit

Training observations and fitted polynomial.

Diagnosing…

Model complexity curve

Shaded zones separate underfit, useful-fit and overfit regions; the validation minimum marks the sweet spot.

Learning curve

The shaded gap between train and validation error shows how much generalisation improves as more training data is added.