Interpretation & Communication · Lesson 75

Permutation Importance

Permutation Importance is part of model interpretation and communication.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What Permutation Importance actually means

Permutation Importance is part of model interpretation and communication. Explanations describe model behaviour under specific assumptions; they are not automatically causal statements about the real world.

Permutation Importance matters because interpretation is useful only after predictive validity is established and only within the assumptions of the explanation method. Communication must distinguish model reliance, statistical association and causal effect.

Deeper walkthrough

Read Permutation Importance as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Start with the audience and decision. Stage 2: Choose an explanation method that matches model type and question (global vs local). Stage 3: Use held-out data when measuring importance that depends on predictive performance. Final checkpoint: State uncertainty, limitations and what the explanation does not establish.

Mechanism

Follow the transformation

Start with the audience and decision.

Choose an explanation method that matches model type and question (global vs local).

Use held-out data when measuring importance that depends on predictive performance.

Evidence

Know what would convince you

  • Repeat the explanation on held-out/subsampled data and check whether the important pattern is stable.
  • Compare at least one alternative explanation or direct prediction perturbation for a small case.
Useful distinctionCoefficients: Direct model parameters; interpretation depends on scaling/encoding and model form.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Start with the audience and decision

Start with the audience and decision. For Permutation Importance, identify the exact state before this stage, the operation or rule applied here, and the observable state afterwards so the mechanism remains inspectable.

State focus: identify exactly what changed at this stage and what observable evidence confirms that change.
How it works

Trace the mechanism step by step

  1. Start with the audience and decision.
  2. Choose an explanation method that matches model type and question (global vs local).
  3. Use held-out data when measuring importance that depends on predictive performance.
  4. Check correlated features because importance can be shared or displaced.
  5. State uncertainty, limitations and what the explanation does not establish.
Worked demonstration

Make the concept concrete

Demonstration

Text example

Baseline validation score = 0.88.
Shuffle feature A → score 0.80 (drop 0.08).
Shuffle feature B → score 0.87 (drop 0.01).
Expected / illustrative result
Feature A is more important to this fitted model on this evaluation set, but correlation can redistribute importance among redundant predictors.
Interpret the result.

For Permutation Importance, connect the displayed result to the specific input and mechanism above; independently verify one value/state change rather than treating successful execution as proof.

Distinctions & related ideas

Know what this is — and what it is not

CoefficientsDirect model parameters; interpretation depends on scaling/encoding and model form.
Permutation importancePerformance loss after shuffling a feature.
PDPAverage predicted response as a feature is varied, marginalising over data.
ICEIndividual prediction curves rather than the average.
SHAPAdditive feature-attribution framework tied to a background/reference distribution.
Use deliberately

When it is appropriate

Use Permutation Importance when the explanation question is explicit—global behaviour, local prediction, feature effect or communication—and the method’s limitations are acceptable.

Boundary conditions

When to stop or reconsider

Do not treat model explanations as causal effects or ground truth, especially with correlated features, extrapolation or unstable models.

Common mistakes

Failure modes to recognise

  • Presenting feature importance or attribution as causality.
  • Ignoring correlated/interacting features that can redistribute or mask apparent importance.
  • Using explanation data outside the region where the fitted model has support and then over-interpreting extrapolation.
Verification

How to check the result

  • Repeat the explanation on held-out/subsampled data and check whether the important pattern is stable.
  • Compare at least one alternative explanation or direct prediction perturbation for a small case.
  • State what the explanation depends on—model, background data, feature dependence and prediction point—before drawing a decision conclusion.
Feature importance & selection workflow

Calculate feature evidence before choosing what to keep

No single feature-importance number answers every question. Use statistical, univariate predictive and model-based evidence together, then validate the selected subset.

Inspect distributions→Welch t-test + p-value→Univariate ROC-AUC→Fit model→Permutation importance→RFE→Validate subset

What each method asks

  • Welch t-test / p-value: is the class-mean difference inconsistent with the equal-means null under the test assumptions?
  • AUC-based feature importance: use univariate max(AUC, 1-AUC) to measure one-feature discrimination, and/or measure the drop in model ROC-AUC when the fitted model loses that feature’s information.
  • Permutation importance: how much does a fitted model’s held-out performance fall when this feature is disrupted?
  • RFE: which features survive repeated model-based elimination, with subset size chosen using validation evidence?

Do not collapse the meanings

Important
A small p-value is not the same as a large predictive effect. A high univariate AUC does not prove the feature is needed in a multivariate model. RFE is model-dependent. None of these methods establishes causality.
Validation rule
Fit screening thresholds, scaling and RFE inside the training/validation workflow. Keep the final test set outside feature selection.
# Feature evidence on training data only
# Step 1 — Import the module so its functions/classes are available to the rest of this example.
import numpy as np
# Step 2 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from scipy.stats import ttest_ind
# Step 3 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.metrics import roc_auc_score
# Step 4 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.linear_model import LogisticRegression
# Step 5 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.feature_selection import RFE
# Step 6 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.inspection import permutation_importance

# Step 7 — Compute the right-hand expression and store its result in `feature_rows` for the next step.
feature_rows = []
# Step 8 — Iterate through the collection so the indented block is applied once for each item.
for name in feature_names:
    # Step 9 — Execute this statement and inspect how it changes the current value, object or program state.
    x0 = X_train.loc[y_train == 0, name].dropna()
    # Step 10 — Execute this statement and inspect how it changes the current value, object or program state.
    x1 = X_train.loc[y_train == 1, name].dropna()

    # 1) Statistical significance: unequal-variance two-sample test
    # Step 11 — Compute the right-hand expression and store its result in `t_stat, p_value` for the next step.
    t_stat, p_value = ttest_ind(x0, x1, equal_var=False)

    # 2) Univariate discrimination; direction-adjusted for importance
    # Step 12 — Compute the right-hand expression and store its result in `auc` for the next step.
    auc = roc_auc_score(y_train, X_train[name])
    # Step 13 — Compute the right-hand expression and store its result in `auc_importance` for the next step.
    auc_importance = max(auc, 1 - auc)

    # Step 14 — Execute this statement and inspect how it changes the current value, object or program state.
    feature_rows.append((name, t_stat, p_value, auc_importance))

# 3) Model-based subset selection: fit only on development data
# Step 15 — Instantiate `base` with the chosen algorithm/configuration before fitting it to data.
base = LogisticRegression(max_iter=2000)
# Step 16 — Compute the right-hand expression and store its result in `rfe` for the next step.
rfe = RFE(base, n_features_to_select=3)
# Step 17 — Fit the model or transformer, learning its parameters from the supplied training data.
rfe.fit(X_train_scaled, y_train)
# Step 18 — Construct `selected` as an array so vectorised numerical operations can be applied consistently.
selected = np.array(feature_names)[rfe.support_]

# 4) Model reliance: evaluate permutation loss on validation data
# Step 19 — Fit the model or transformer, learning its parameters from the supplied training data.
model = LogisticRegression(max_iter=2000).fit(
    X_train_scaled[:, rfe.support_], y_train
)
# Step 20 — Compute the right-hand expression and store its result in `perm` for the next step.
perm = permutation_importance(
    model,
    X_valid_scaled[:, rfe.support_],
    y_valid,
    scoring="roc_auc",
    n_repeats=20,
    random_state=42,
)

# Step 21 — Display the current value explicitly so the result/state can be inspected during execution.
print("Selected by RFE:", selected.tolist())
# Step 22 — Display the current value explicitly so the result/state can be inspected during execution.
print("Validation permutation ΔAUC:", perm.importances_mean.round(3))
Multiple features mean multiple tests

If many p-values are screened at once, consider false-discovery or family-wise error control and report effect sizes as well as significance. Use the p-value as evidence under assumptions, not as a mechanical feature-selection cutoff.

Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of Permutation Importance. First start with the audience and decision. Then choose an explanation method that matches model type and question (global vs local). Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Pick one prediction or a small held-out sample. Change one feature/condition deliberately and compare the model response with the explanation you expected.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from Permutation Importance, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Start with the audience and decision.
Step 2Choose an explanation method that matches model type and question (global vs local).
Step 3Use held-out data when measuring importance that depends on predictive performance.
Step 4Check correlated features because importance can be shared or displaced.
Lesson summary

What to remember

  • Permutation Importance is part of model interpretation and communication. Explanations describe model behaviour under specific assumptions; they are not automatically causal statements about the real world.
  • Start with the audience and decision.
  • Presenting feature importance or attribution as causality.
  • Repeat the explanation on held-out/subsampled data and check whether the important pattern is stable.