Tree Models · Lesson 41

Feature Importance Caveats

Feature Importance Caveats is part of decision-tree learning.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What Feature Importance Caveats actually means

Feature Importance Caveats is part of decision-tree learning. A tree recursively partitions feature space using threshold/category splits chosen to improve node purity or reduce prediction error; terminal leaves store class distributions or numeric predictions.

Feature Importance Caveats matters because tree models learn nonlinear threshold rules and interactions with little preprocessing, but unconstrained trees can have high variance. Ensembles and pruning trade interpretability, variance and computational cost in different ways.

Deeper walkthrough

Read Feature Importance Caveats as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Start with all training samples at the root. Stage 2: Evaluate candidate feature splits. Stage 3: Choose a split that most reduces impurity/error. Final checkpoint: Predict by following one path from root to leaf.

Mechanism

Follow the transformation

Start with all training samples at the root.

Evaluate candidate feature splits.

Choose a split that most reduces impurity/error.

Evidence

Know what would convince you

  • Fit a tiny or baseline case first and confirm prediction shape/range and a few outputs.
  • Evaluate with the same held-out folds/metric as competing models and inspect variability, not just the mean.
Useful distinctionClassification tree: Leaves predict classes/probabilities; split criteria include Gini/entropy.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Start with all training samples at…

Start with all training samples at the root. At this stage of Feature Importance Caveats, keep the incoming data or object separate from the learned parameter, transformed object, or statistic so the change can be reproduced and independently checked.

Transformation focus: keep the input and produced parameters/result separate so the change is observable and reproducible.
How it works

Trace the mechanism step by step

  1. Start with all training samples at the root.
  2. Evaluate candidate feature splits.
  3. Choose a split that most reduces impurity/error.
  4. Repeat recursively in child nodes.
  5. Stop or prune using depth, minimum samples, cost-complexity or validation criteria.
  6. Predict by following one path from root to leaf.
Worked demonstration

Make the concept concrete

Demonstration

Text example

Correlated features A and B carry nearly the same signal.
A tree may split mostly on A, assigning low impurity importance to B.
Permutation of A alone may show modest loss because B can substitute.
Expected / illustrative result
Importance is conditional on model, data and correlated alternatives; it is not a causal ranking of real-world factors.
Interpret the result.

For Feature Importance Caveats, connect the reported result to the exact training/validation/prediction step that produced it and check one prediction, fold or metric component independently.

Distinctions & related ideas

Know what this is — and what it is not

Classification treeLeaves predict classes/probabilities; split criteria include Gini/entropy.
Regression treeLeaves predict numeric values; splits reduce squared/absolute error.
Deep treeLow bias, high variance; can memorise training details.
Pruned/shallow treeHigher bias, often better stability/generalisation.
Use deliberately

When it is appropriate

Use Feature Importance Caveats when its inductive assumptions fit the feature/target structure and it can be compared fairly with a simpler baseline on unseen data.

Boundary conditions

When to stop or reconsider

Prefer a simpler or different model when the sample size, representation, computational budget, interpretability requirement or data geometry conflicts with this method.

Common mistakes

Failure modes to recognise

  • Judging the model only by training fit instead of generalisation on held-out data.
  • Comparing models with inconsistent preprocessing, folds or evaluation metrics.
  • Tuning complexity without checking a simple baseline, error patterns and variance across splits.
Verification

How to check the result

  • Fit a tiny or baseline case first and confirm prediction shape/range and a few outputs.
  • Evaluate with the same held-out folds/metric as competing models and inspect variability, not just the mean.
  • Inspect errors/residuals or decision boundaries and vary one key hyperparameter to verify expected behaviour.
Feature importance & selection workflow

Calculate feature evidence before choosing what to keep

No single feature-importance number answers every question. Use statistical, univariate predictive and model-based evidence together, then validate the selected subset.

Inspect distributions→Welch t-test + p-value→Univariate ROC-AUC→Fit model→Permutation importance→RFE→Validate subset

What each method asks

  • Welch t-test / p-value: is the class-mean difference inconsistent with the equal-means null under the test assumptions?
  • AUC-based feature importance: use univariate max(AUC, 1-AUC) to measure one-feature discrimination, and/or measure the drop in model ROC-AUC when the fitted model loses that feature’s information.
  • Permutation importance: how much does a fitted model’s held-out performance fall when this feature is disrupted?
  • RFE: which features survive repeated model-based elimination, with subset size chosen using validation evidence?

Do not collapse the meanings

Important
A small p-value is not the same as a large predictive effect. A high univariate AUC does not prove the feature is needed in a multivariate model. RFE is model-dependent. None of these methods establishes causality.
Validation rule
Fit screening thresholds, scaling and RFE inside the training/validation workflow. Keep the final test set outside feature selection.
# Feature evidence on training data only
# Step 1 — Import the module so its functions/classes are available to the rest of this example.
import numpy as np
# Step 2 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from scipy.stats import ttest_ind
# Step 3 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.metrics import roc_auc_score
# Step 4 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.linear_model import LogisticRegression
# Step 5 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.feature_selection import RFE
# Step 6 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.inspection import permutation_importance

# Step 7 — Compute the right-hand expression and store its result in `feature_rows` for the next step.
feature_rows = []
# Step 8 — Iterate through the collection so the indented block is applied once for each item.
for name in feature_names:
    # Step 9 — Execute this statement and inspect how it changes the current value, object or program state.
    x0 = X_train.loc[y_train == 0, name].dropna()
    # Step 10 — Execute this statement and inspect how it changes the current value, object or program state.
    x1 = X_train.loc[y_train == 1, name].dropna()

    # 1) Statistical significance: unequal-variance two-sample test
    # Step 11 — Compute the right-hand expression and store its result in `t_stat, p_value` for the next step.
    t_stat, p_value = ttest_ind(x0, x1, equal_var=False)

    # 2) Univariate discrimination; direction-adjusted for importance
    # Step 12 — Compute the right-hand expression and store its result in `auc` for the next step.
    auc = roc_auc_score(y_train, X_train[name])
    # Step 13 — Compute the right-hand expression and store its result in `auc_importance` for the next step.
    auc_importance = max(auc, 1 - auc)

    # Step 14 — Execute this statement and inspect how it changes the current value, object or program state.
    feature_rows.append((name, t_stat, p_value, auc_importance))

# 3) Model-based subset selection: fit only on development data
# Step 15 — Instantiate `base` with the chosen algorithm/configuration before fitting it to data.
base = LogisticRegression(max_iter=2000)
# Step 16 — Compute the right-hand expression and store its result in `rfe` for the next step.
rfe = RFE(base, n_features_to_select=3)
# Step 17 — Fit the model or transformer, learning its parameters from the supplied training data.
rfe.fit(X_train_scaled, y_train)
# Step 18 — Construct `selected` as an array so vectorised numerical operations can be applied consistently.
selected = np.array(feature_names)[rfe.support_]

# 4) Model reliance: evaluate permutation loss on validation data
# Step 19 — Fit the model or transformer, learning its parameters from the supplied training data.
model = LogisticRegression(max_iter=2000).fit(
    X_train_scaled[:, rfe.support_], y_train
)
# Step 20 — Compute the right-hand expression and store its result in `perm` for the next step.
perm = permutation_importance(
    model,
    X_valid_scaled[:, rfe.support_],
    y_valid,
    scoring="roc_auc",
    n_repeats=20,
    random_state=42,
)

# Step 21 — Display the current value explicitly so the result/state can be inspected during execution.
print("Selected by RFE:", selected.tolist())
# Step 22 — Display the current value explicitly so the result/state can be inspected during execution.
print("Validation permutation ΔAUC:", perm.importances_mean.round(3))
Multiple features mean multiple tests

If many p-values are screened at once, consider false-discovery or family-wise error control and report effect sizes as well as significance. Use the p-value as evidence under assumptions, not as a mechanical feature-selection cutoff.

Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of Feature Importance Caveats. First start with all training samples at the root. Then evaluate candidate feature splits. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Start with a small baseline and a fixed validation split/fold assignment. Predict what increasing or decreasing one complexity control should do before testing it.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from Feature Importance Caveats, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Start with all training samples at the root.
Step 2Evaluate candidate feature splits.
Step 3Choose a split that most reduces impurity/error.
Step 4Repeat recursively in child nodes.
Lesson summary

What to remember

  • Feature Importance Caveats is part of decision-tree learning. A tree recursively partitions feature space using threshold/category splits chosen to improve node purity or reduce prediction error; terminal leaves store class distributions or numeric predictions.
  • Start with all training samples at the root.
  • Judging the model only by training fit instead of generalisation on held-out data.
  • Fit a tiny or baseline case first and confirm prediction shape/range and a few outputs.