Math for ML · Lesson 10

Loss Functions

A loss function converts prediction error into a numerical penalty that the learning algorithm can optimise.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

Loss Functions

A loss function converts prediction error into a numerical penalty that the learning algorithm can optimise. Squared error heavily penalises large regression residuals; absolute error grows linearly; log loss evaluates probabilistic classification and strongly penalises confident wrong probabilities. The loss defines what “fit the training data” mathematically means.

Learning goal: explain why Loss Functions behaves this way, apply it to a small example, and verify the result independently. Begin by being able to justify this first step: Generate a prediction from the current model.

Deeper walkthrough

Read Loss Functions as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Generate a prediction from the current model. Stage 2: Compare it with the target using the chosen loss. Stage 3: Aggregate losses across training examples into an objective, optionally adding regularisation. Final checkpoint: Evaluate final performance with decision-relevant held-out metrics, which need not be identical to the training loss.

Mechanism

Follow the transformation

Generate a prediction from the current model.

Compare it with the target using the chosen loss.

Aggregate losses across training examples into an objective, optionally adding regularisation.

Evidence

Know what would convince you

  • Verify the split/validation boundary before comparing scores.
  • Inspect model/preprocessing state or a hand-computable tiny example.
Useful distinctionRepresentation: How the method encodes inputs/predictions.
Visual demonstration of Loss Functions
Visual demonstration: use the diagram to trace the main objects and state changes involved in Loss Functions.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Generate a prediction from the current…

Generate a prediction from the current model. Treat the output from Loss Functions as evidence to inspect: confirm its type, shape, range or units and connect it back to the input that produced it.

Output focus: inspect both the value and its shape/type/meaning before treating it as a trustworthy result.
How it works

Trace the mechanism step by step

  1. Generate a prediction from the current model.
  2. Compare it with the target using the chosen loss.
  3. Aggregate losses across training examples into an objective, optionally adding regularisation.
  4. Optimise model parameters to reduce that objective.
  5. Evaluate final performance with decision-relevant held-out metrics, which need not be identical to the training loss.
Worked demonstration

Compare regression losses

Target y=10
Prediction A=9 → residual=1 → squared loss=1
Prediction B=6 → residual=4 → squared loss=16
Expected / illustrative result
Squared error gives the four-unit miss sixteen times the penalty of the one-unit miss.
Interpret the result.

For Loss Functions, identify exactly what each reported quantity represents, including its units/denominator, and independently recompute one part of the result.

Distinctions & related ideas

Place the concept correctly

RepresentationHow the method encodes inputs/predictions.
Learning/operationWhat fitted state or calculation changes.
ValidationIndependent evidence used to judge generalisation or correctness.
Use deliberately

When it is appropriate

Use Loss Functions when it answers a defined question in Math for ML and its inputs/assumptions match the current data or program state.

Boundary conditions

When to stop or reconsider

Reconsider Loss Functions when the required information is unavailable, the operation would violate a validation/data boundary, or a simpler operation answers the question more transparently.

Common mistakes

Failure modes to recognise

  • Optimising on the final test set.
  • Ignoring feature scale/representation or split structure when the method depends on them.
  • Reporting a single score without checking errors, variance or operating conditions.
Verification

How to check the result

  • Verify the split/validation boundary before comparing scores.
  • Inspect model/preprocessing state or a hand-computable tiny example.
  • Perturb one input/hyperparameter and predict the expected direction or behaviour.
Hands-on practice

Demonstrate understanding

Try this:

Construct a tiny example of Loss Functions. First generate a prediction from the current model. Then compare it with the target using the chosen loss. Predict the result before execution and explain one boundary or failure case.

Use a very small example and calculate one quantity manually. Separate sample evidence from population/causal claims.
Knowledge check

Check reasoning, not memorisation

Which approach best demonstrates understanding of Loss Functions?

Quick reference

Remember the logic

Step 1Generate a prediction from the current model.
Step 2Compare it with the target using the chosen loss.
Step 3Aggregate losses across training examples into an objective, optionally adding regularisation.
Step 4Optimise model parameters to reduce that objective.
Lesson summary

What to remember

  • A loss function converts prediction error into a numerical penalty that the learning algorithm can optimise. Squared error heavily penalises large regression residuals; absolute error grows linearly; log loss evaluates probabilistic classification and strongly penalises confident wrong probabilities. The loss defines what “fit the training data” mathematically means.
  • Generate a prediction from the current model.
  • Optimising on the final test set.
  • Verify the split/validation boundary before comparing scores.