Interpretation · Flagship experience

Explainability & Model Interpretation

What kind of explanation do you actually need?

Start here

What kind of explanation do you actually need?

Feature evidence can come from statistical significance, single-feature discrimination, model reliance, subset selection or local prediction effects. These are related but not interchangeable.

Building interactive view…
Feature evidence workflow

Calculate importance with more than one lens.

Statistical significance, single-feature discrimination, model reliance and recursive selection answer different questions. Use them together rather than treating one ranking as truth.

Open Explainable AI Playground →
Inspect→Welch t-test + p-value→Univariate ROC-AUC→Fit model→Permutation importance→RFE→Local explanation

p-value evidence

Welch’s t-test compares the class means for each numeric feature.

Univariate feature AUC

Direction-adjusted ROC-AUC shows how well each feature alone ranks the classes.

Permutation importance

Shuffle one feature and measure held-out ROC-AUC and F1 loss.

Recursive Feature Elimination

Repeatedly remove the weakest standardized logistic feature and inspect validation F1.

Interpretation rule: a small p-value does not prove predictive usefulness; a high AUC does not prove multivariate necessity; a high permutation score does not prove causality; and an RFE rank is model- and validation-dependent.
Understand

Build the mental model

Feature evidence can come from statistical significance, single-feature discrimination, model reliance, subset selection or local prediction effects. These are related but not interchangeable. Use multiple evidence types, keep selection inside training/validation data, account for multiple testing where appropriate, and never equate importance with causality.

What happens if…?

Break the assumption deliberately

Duplicate a correlated feature and observe how importance can split or move between substitutes.

Move the control and explain what you expect before reading the visual.

Technical lens

Formalise what the visual is doing

Welch t-tests produce p-values for class-mean differences; univariate ROC-AUC measures ranking discrimination; permutation importance measures held-out performance loss; RFE recursively removes weak model features; local methods explain individual predictions.

Technical questionUse a tiny case to make the mechanism observable. Welch t-tests produce p-values for class-mean differences; univariate ROC-AUC measures ranking discrimination; permutation importance measures held-out performance loss; RFE recursively removes weak model features; local methods explain individual predictions. Verify one intermediate quantity, state change or mapping independently; then predict the consequence of this change: Duplicate a correlated feature and observe how importance can split or move between substitutes.
Practitioner lens

Use it responsibly

Use multiple evidence types, keep selection inside training/validation data, account for multiple testing where appropriate, and never equate importance with causality.

Transfer testTransfer this idea to a new example and justify each decision using this practitioner rule: Use multiple evidence types, keep selection inside training/validation data, account for multiple testing where appropriate, and never equate importance with causality. Then explain what should change if you deliberately test: Duplicate a correlated feature and observe how importance can split or move between substitutes.
Worked exploration

Use the visual as an experiment, not decoration

For one fitted model, compare coefficient magnitude, permutation importance and a local SHAP-style attribution. Each answers a different question. Introduce two correlated features and observe how importance can be shared or displaced.

Technical lens

Welch t-tests produce p-values for class-mean differences; univariate ROC-AUC measures ranking discrimination; permutation importance measures held-out performance loss; RFE recursively removes weak model features; local methods explain individual predictions.

Practitioner check

Use multiple evidence types, keep selection inside training/validation data, account for multiple testing where appropriate, and never equate importance with causality.

Prediction before interaction
Duplicate a correlated feature and observe how importance can split or move between substitutes.
Exploration walkthrough

Turn the interaction into an evidence trail

For one fitted model, compare coefficient magnitude, permutation importance and a local SHAP-style attribution. Each answers a different question. Introduce two correlated features and observe how importance can be shared or displaced. Before moving the control, state your prediction. After the visual changes, name the specific state, statistic, boundary or mapping that changed and explain why that change is consistent—or inconsistent—with your prediction.

  • Record one observable quantity before the interaction and the same quantity afterwards.
  • Change one factor at a time so the causal effect of the control is inspectable.
  • Use an edge or failure case to discover where the concept stops behaving as the simple story suggests.
Reference depth

Open the complete material

The flagship experience is the map. These pages contain the roads.

Continue this exact concept

Choose depth, practice or application.

These destinations are explicitly mapped to Explainability & Model Interpretation; they are not generic landing-page fallbacks.