Follow the transformation
Generate a prediction from the current model.
Compare it with the target using the chosen loss.
Aggregate losses across training examples into an objective, optionally adding regularisation.
A loss function converts prediction error into a numerical penalty that the learning algorithm can optimise.
A loss function converts prediction error into a numerical penalty that the learning algorithm can optimise. Squared error heavily penalises large regression residuals; absolute error grows linearly; log loss evaluates probabilistic classification and strongly penalises confident wrong probabilities. The loss defines what “fit the training data” mathematically means.
Learning goal: explain why Loss Functions behaves this way, apply it to a small example, and verify the result independently. Begin by being able to justify this first step: Generate a prediction from the current model.
Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Generate a prediction from the current model. Stage 2: Compare it with the target using the chosen loss. Stage 3: Aggregate losses across training examples into an objective, optionally adding regularisation. Final checkpoint: Evaluate final performance with decision-relevant held-out metrics, which need not be identical to the training loss.
Generate a prediction from the current model.
Compare it with the target using the chosen loss.
Aggregate losses across training examples into an objective, optionally adding regularisation.
Generate a prediction from the current model. Treat the output from Loss Functions as evidence to inspect: confirm its type, shape, range or units and connect it back to the input that produced it.
Target y=10
Prediction A=9 → residual=1 → squared loss=1
Prediction B=6 → residual=4 → squared loss=16Squared error gives the four-unit miss sixteen times the penalty of the one-unit miss.
For Loss Functions, identify exactly what each reported quantity represents, including its units/denominator, and independently recompute one part of the result.
RepresentationHow the method encodes inputs/predictions.Learning/operationWhat fitted state or calculation changes.ValidationIndependent evidence used to judge generalisation or correctness.Use Loss Functions when it answers a defined question in Math for ML and its inputs/assumptions match the current data or program state.
Reconsider Loss Functions when the required information is unavailable, the operation would violate a validation/data boundary, or a simpler operation answers the question more transparently.
Construct a tiny example of Loss Functions. First generate a prediction from the current model. Then compare it with the target using the chosen loss. Predict the result before execution and explain one boundary or failure case.
Which approach best demonstrates understanding of Loss Functions?
Step 1Generate a prediction from the current model.Step 2Compare it with the target using the chosen loss.Step 3Aggregate losses across training examples into an objective, optionally adding regularisation.Step 4Optimise model parameters to reduce that objective.