Predictive Modelling · Lesson 62

Random Forest

A random forest averages many decision trees built from resampled observations and randomly restricted feature candidates.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

Random Forest

A random forest averages many decision trees built from resampled observations and randomly restricted feature candidates. The diversity between trees reduces the variance of a single tree while retaining nonlinear interactions and threshold behaviour.

Learning goal: explain why Random Forest behaves this way, apply it to a small example, and verify the result independently. Begin by being able to justify this first step: Draw bootstrap samples for individual trees.

Deeper walkthrough

Read Random Forest as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Draw bootstrap samples for individual trees. Stage 2: At each split, consider a random subset of features. Stage 3: Grow many trees with controlled hyperparameters. Final checkpoint: Use out-of-bag/validation evidence and interpret impurity importances cautiously.

Mechanism

Follow the transformation

Draw bootstrap samples for individual trees.

At each split, consider a random subset of features.

Grow many trees with controlled hyperparameters.

Evidence

Know what would convince you

  • Confirm fitted transformations/models saw only training data.
  • Retain fold/test predictions so metrics can be recomputed independently.
Useful distinctionTraining evidence: Information allowed to influence fitted state.
Visual demonstration of Random Forest
Visual demonstration: use the diagram to trace the main objects and state changes involved in Random Forest.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Draw bootstrap samples for individual trees

Draw bootstrap samples for individual trees. For Random Forest, identify the exact state before this stage, the operation or rule applied here, and the observable state afterwards so the mechanism remains inspectable.

State focus: identify exactly what changed at this stage and what observable evidence confirms that change.
How it works

Trace the mechanism step by step

  1. Draw bootstrap samples for individual trees.
  2. At each split, consider a random subset of features.
  3. Grow many trees with controlled hyperparameters.
  4. Average predictions or majority-vote classes.
  5. Use out-of-bag/validation evidence and interpret impurity importances cautiously.
Worked demonstration

Ensemble vote

Tree votes: [A, A, B, A, B] → forest prediction A.
Expected / illustrative result
No single tree controls the result; aggregation stabilises predictions across varied training samples/features.
Interpret the result.

For Random Forest, connect the result to the fitted state, held-out data or prediction rule that produced it and independently check one prediction, split or metric component.

Distinctions & related ideas

Place the concept correctly

Training evidenceInformation allowed to influence fitted state.
Held-out evidenceIndependent observations used to estimate generalisation.
InterpretationWhat the result supports, with assumptions and limitations.
Use deliberately

When it is appropriate

Use Random Forest when it answers a defined question in Predictive Modelling and its inputs/assumptions match the current data or program state.

Boundary conditions

When to stop or reconsider

Reconsider Random Forest when the required information is unavailable, the operation would violate a validation/data boundary, or a simpler operation answers the question more transparently.

Common mistakes

Failure modes to recognise

  • Learning preprocessing/feature/model choices from held-out test information.
  • Comparing models under different splits or preprocessing and attributing the difference to the algorithm.
  • Turning an association or model explanation into an unsupported causal claim.
Verification

How to check the result

  • Confirm fitted transformations/models saw only training data.
  • Retain fold/test predictions so metrics can be recomputed independently.
  • Inspect errors/subgroups and compare with a baseline before generalising the conclusion.
Hands-on practice

Demonstrate understanding

Try this:

Construct a tiny example of Random Forest. First draw bootstrap samples for individual trees. Then at each split, consider a random subset of features. Predict the result before execution and explain one boundary or failure case.

Use a tiny fixed split or synthetic example. State what is fitted, what remains held out, and what result you expect before running it.
Knowledge check

Check reasoning, not memorisation

Which approach best demonstrates understanding of Random Forest?

Quick reference

Remember the logic

Step 1Draw bootstrap samples for individual trees.
Step 2At each split, consider a random subset of features.
Step 3Grow many trees with controlled hyperparameters.
Step 4Average predictions or majority-vote classes.
Lesson summary

What to remember

  • A random forest averages many decision trees built from resampled observations and randomly restricted feature candidates. The diversity between trees reduces the variance of a single tree while retaining nonlinear interactions and threshold behaviour.
  • Draw bootstrap samples for individual trees.
  • Learning preprocessing/feature/model choices from held-out test information.
  • Confirm fitted transformations/models saw only training data.