Predictive Modelling · Lesson 63

Gradient Boosting

Gradient boosting builds an additive model sequentially.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

Gradient Boosting

Gradient boosting builds an additive model sequentially. Each new weak learner, usually a shallow tree, is fitted to reduce the errors left by the current ensemble—more generally, to follow the negative gradient of the chosen loss.

Learning goal: explain why Gradient Boosting behaves this way, apply it to a small example, and verify the result independently. Begin by being able to justify this first step: Start from a simple initial prediction.

Deeper walkthrough

Read Gradient Boosting as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Start from a simple initial prediction. Stage 2: Compute residual-like gradients under the loss. Stage 3: Fit a weak learner to that remaining signal. Final checkpoint: Repeat while monitoring validation performance/early stopping.

Mechanism

Follow the transformation

Start from a simple initial prediction.

Compute residual-like gradients under the loss.

Fit a weak learner to that remaining signal.

Evidence

Know what would convince you

  • Confirm fitted transformations/models saw only training data.
  • Retain fold/test predictions so metrics can be recomputed independently.
Useful distinctionTraining evidence: Information allowed to influence fitted state.
Visual demonstration of Gradient Boosting
Visual demonstration: use the diagram to trace the main objects and state changes involved in Gradient Boosting.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Start from a simple initial prediction

Start from a simple initial prediction. Treat the output from Gradient Boosting as evidence to inspect: confirm its type, shape, range or units and connect it back to the input that produced it.

Output focus: inspect both the value and its shape/type/meaning before treating it as a trustworthy result.
How it works

Trace the mechanism step by step

  1. Start from a simple initial prediction.
  2. Compute residual-like gradients under the loss.
  3. Fit a weak learner to that remaining signal.
  4. Add its prediction scaled by a learning rate.
  5. Repeat while monitoring validation performance/early stopping.
Worked demonstration

Sequential correction

Model 0 predicts 10, target is 14 → residual +4.
A new tree predicts a +3 correction (after learning-rate scaling) → ensemble moves toward 13.
Expected / illustrative result
Boosting improves the ensemble stage by stage rather than averaging independent trees.
Interpret the result.

For Gradient Boosting, connect the result to the fitted state, held-out data or prediction rule that produced it and independently check one prediction, split or metric component.

Distinctions & related ideas

Place the concept correctly

Training evidenceInformation allowed to influence fitted state.
Held-out evidenceIndependent observations used to estimate generalisation.
InterpretationWhat the result supports, with assumptions and limitations.
Use deliberately

When it is appropriate

Use Gradient Boosting when it answers a defined question in Predictive Modelling and its inputs/assumptions match the current data or program state.

Boundary conditions

When to stop or reconsider

Reconsider Gradient Boosting when the required information is unavailable, the operation would violate a validation/data boundary, or a simpler operation answers the question more transparently.

Common mistakes

Failure modes to recognise

  • Learning preprocessing/feature/model choices from held-out test information.
  • Comparing models under different splits or preprocessing and attributing the difference to the algorithm.
  • Turning an association or model explanation into an unsupported causal claim.
Verification

How to check the result

  • Confirm fitted transformations/models saw only training data.
  • Retain fold/test predictions so metrics can be recomputed independently.
  • Inspect errors/subgroups and compare with a baseline before generalising the conclusion.
Hands-on practice

Demonstrate understanding

Try this:

Construct a tiny example of Gradient Boosting. First start from a simple initial prediction. Then compute residual-like gradients under the loss. Predict the result before execution and explain one boundary or failure case.

Use a tiny fixed split or synthetic example. State what is fitted, what remains held out, and what result you expect before running it.
Knowledge check

Check reasoning, not memorisation

Which approach best demonstrates understanding of Gradient Boosting?

Quick reference

Remember the logic

Step 1Start from a simple initial prediction.
Step 2Compute residual-like gradients under the loss.
Step 3Fit a weak learner to that remaining signal.
Step 4Add its prediction scaled by a learning rate.
Lesson summary

What to remember

  • Gradient boosting builds an additive model sequentially. Each new weak learner, usually a shallow tree, is fitted to reduce the errors left by the current ensemble—more generally, to follow the negative gradient of the chosen loss.
  • Start from a simple initial prediction.
  • Learning preprocessing/feature/model choices from held-out test information.
  • Confirm fitted transformations/models saw only training data.