Evaluation, Robustness & Advanced ML

Explainable AI Playground

Trace feature evidence from statistical screening and univariate discrimination through model-based importance, RFE and local prediction explanation.

Lab concept guide

What to observe while you experiment

Model explanation methods answer different questions: screening describes marginal association, model-based importance reflects fitted structure, permutation measures performance sensitivity, RFE evaluates subsets, and local explanations attribute one prediction. None automatically implies causality.

MechanismFollow a feature from raw association through fitted-model importance to a local prediction contribution and note where the meaning changes.
Failure modeTreating importance/attribution as a causal effect or ignoring correlated features that can share/displace importance.
VerificationPerturb one feature or repeat importance on held-out/resampled data and check whether the claimed explanation is stable and prediction-linked.
Experiment deliberately
Choose one feature, compare its screening score, global importance and local effect, then explain why the three numbers need not rank it identically.
Importance is not one number. Statistical significance, univariate AUC, model reliance, subset selection and local effects answer different questions. Compare them before deciding that a feature “matters”.
Feature-importance calculation process

From evidence to selected features

Use training/validation evidence for screening and selection; reserve held-out evidence for final model assessment.

Explore the concept →
1 · Inspect feature distributions→2 · Welch t-test + p-value→3 · Univariate ROC-AUC→4 · Fit model→5 · Permutation / coefficients→6 · RFE→7 · Local explanation
p-valueEvidence against equal class means under the test assumptions; not predictive importance.
Feature AUCHow well one feature alone ranks positive above negative cases; direction-adjusted AUC is used for importance.
RFERepeatedly removes the weakest model feature and validates the remaining subset.
Ready.

Feature evidence table

Compare inferential evidence, univariate discrimination, model reliance and RFE rank side by side.

p < 0.05 is a screening flag, not a guarantee

Statistical significance and p-values

Welch’s two-sample t-test compares class means without assuming equal variance.

Univariate AUC feature importance

Each feature is used alone as a ranking score. Values are direction-adjusted so 0.5 = no discrimination and 1.0 = perfect separation.

Recursive Feature Elimination (RFE)

Standardised logistic coefficients determine the next feature removed; validation F1 shows how the remaining subset behaves.

RFE elimination path

Global permutation importance

Performance drop after shuffling one feature at a time, shown for both ROC-AUC and F1.

Local feature effects

Each feature is replaced by its training median to form a simple counterfactual.

Model separator and explained observation