Follow the transformation
Evaluate candidate splits on training data.
Choose the split with the best impurity/error improvement.
Repeat recursively in child nodes.
A decision tree predicts by recursively splitting the feature space into regions.
A decision tree predicts by recursively splitting the feature space into regions. At each node it chooses a feature and threshold/category rule that reduces impurity for classification or prediction error for regression, then continues until a stopping or pruning condition is reached.
Learning goal: explain why Decision Trees behaves this way, apply it to a small example, and verify the result independently. Begin by being able to justify this first step: Evaluate candidate splits on training data.
Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Evaluate candidate splits on training data. Stage 2: Choose the split with the best impurity/error improvement. Stage 3: Repeat recursively in child nodes. Final checkpoint: Validate because deep trees can fit noise very closely.
Evaluate candidate splits on training data.
Choose the split with the best impurity/error improvement.
Repeat recursively in child nodes.
Evaluate candidate splits on training data. At this stage of Decision Trees, keep the incoming data or object separate from the learned parameter, transformed object, or statistic so the change can be reproduced and independently checked.
If age < 30 → class A; otherwise → class B.A tree is a sequence of explicit decision rules; a real fitted tree chooses thresholds from data.
For Decision Trees, connect the result to the fitted state, held-out data or prediction rule that produced it and independently check one prediction, split or metric component.
Training evidenceInformation allowed to influence fitted state.Held-out evidenceIndependent observations used to estimate generalisation.InterpretationWhat the result supports, with assumptions and limitations.Use Decision Trees when it answers a defined question in Predictive Modelling and its inputs/assumptions match the current data or program state.
Reconsider Decision Trees when the required information is unavailable, the operation would violate a validation/data boundary, or a simpler operation answers the question more transparently.
Construct a tiny example of Decision Trees. First evaluate candidate splits on training data. Then choose the split with the best impurity/error improvement. Predict the result before execution and explain one boundary or failure case.
Which approach best demonstrates understanding of Decision Trees?
Step 1Evaluate candidate splits on training data.Step 2Choose the split with the best impurity/error improvement.Step 3Repeat recursively in child nodes.Step 4Control depth/minimum samples or prune to limit variance.