Math for ML · Lesson 9

Probability for Prediction

Probability for Prediction supplies mathematical intuition for how models represent predictions and learn from error.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What Probability for Prediction actually means

Probability for Prediction supplies mathematical intuition for how models represent predictions and learn from error. Vectors encode features/parameters, dot products create linear scores, loss functions quantify disagreement with targets, and gradients point toward local directions of change in differentiable objectives.

Probability for Prediction matters because model predictions and optimisation are expressed through vectors, linear algebra, probability and loss functions. Connecting the mathematics to a concrete prediction or parameter update makes later algorithms much easier to understand.

Deeper walkthrough

Read Probability for Prediction as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Represent inputs/parameters as vectors or matrices. Stage 2: Compute the model score/prediction. Stage 3: Evaluate a loss against the observed target. Final checkpoint: Update parameters in a direction that reduces the objective, balancing fit and generalisation.

Mechanism

Follow the transformation

Represent inputs/parameters as vectors or matrices.

Compute the model score/prediction.

Evaluate a loss against the observed target.

Evidence

Know what would convince you

  • Write the dimensions and formula before substituting values.
  • Calculate a two- or three-value example by hand and compare every intermediate term.
Useful distinctionLoss: Objective contribution for prediction error.
Visual demonstration of Probability for Prediction
Visual demonstration: use the diagram to trace the main objects and state changes involved in Probability for Prediction.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Represent inputs/parameters as vectors or matrices

Represent inputs/parameters as vectors or matrices. For Probability for Prediction, identify the exact state before this stage, the operation or rule applied here, and the observable state afterwards so the mechanism remains inspectable.

State focus: identify exactly what changed at this stage and what observable evidence confirms that change.
How it works

Trace the mechanism step by step

  1. Represent inputs/parameters as vectors or matrices.
  2. Compute the model score/prediction.
  3. Evaluate a loss against the observed target.
  4. Compute or approximate how the loss changes with parameters.
  5. Update parameters in a direction that reduces the objective, balancing fit and generalisation.
Worked demonstration

Make the concept concrete

Demonstration

Python example

# Step 1 — Compute the right-hand expression and store its result in `p` for the next step.
p=0.82
# Step 2 — Iterate through the collection so the indented block is applied once for each item.
for threshold in [0.5,0.9]:
    # Step 3 — Display the current value explicitly so the result/state can be inspected during execution.
    print(threshold, "positive" if p>=threshold else "negative")
Expected / illustrative result
At threshold 0.5 the case is positive; at 0.9 it is negative. Probability estimation and decision threshold are separate steps.
Interpret the result.

For Probability for Prediction, connect the displayed result to the specific input and mechanism above; independently verify one value/state change rather than treating successful execution as proof.

Distinctions & related ideas

Know what this is — and what it is not

LossObjective contribution for prediction error.
GradientVector of partial derivatives of loss with respect to parameters.
Learning rateStep-size hyperparameter for gradient updates.
BiasSystematic error from restrictive assumptions.
VarianceSensitivity of the fitted model to training-sample variation.
Use deliberately

When it is appropriate

Use Probability for Prediction when the mathematical object directly represents the model quantity or transformation being reasoned about and dimensions/units are explicit.

Boundary conditions

When to stop or reconsider

Stop if vectors/matrices, dimensions, signs, scales or probability assumptions are not defined; symbolic manipulation without those definitions can produce a formally valid but meaningless result.

Common mistakes

Failure modes to recognise

  • Combining quantities with incompatible dimensions/shapes or forgetting an intercept/normalisation term.
  • Interpreting coefficient or distance magnitude without considering feature scale.
  • Skipping a hand-computable case and therefore missing sign, axis or indexing errors.
Verification

How to check the result

  • Write the dimensions and formula before substituting values.
  • Calculate a two- or three-value example by hand and compare every intermediate term.
  • Perturb one input/parameter and predict the sign/direction of the result before recomputing it.
Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of Probability for Prediction. First represent inputs/parameters as vectors or matrices. Then compute the model score/prediction. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Use two-dimensional vectors or a handful of probabilities so every arithmetic step fits on paper. Track shape, sign and units explicitly.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from Probability for Prediction, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Represent inputs/parameters as vectors or matrices.
Step 2Compute the model score/prediction.
Step 3Evaluate a loss against the observed target.
Step 4Compute or approximate how the loss changes with parameters.
Lesson summary

What to remember

  • Probability for Prediction supplies mathematical intuition for how models represent predictions and learn from error. Vectors encode features/parameters, dot products create linear scores, loss functions quantify disagreement with targets, and gradients point toward local directions of change in differentiable objectives.
  • Represent inputs/parameters as vectors or matrices.
  • Combining quantities with incompatible dimensions/shapes or forgetting an intercept/normalisation term.
  • Write the dimensions and formula before substituting values.