Correlation Analysis & Feature Selection
Measure linear or monotonic association, inspect a correlation matrix, and select useful non-redundant features using train-side correlation evidence.
What to observe while you experiment
Correlation screening combines two distinct checks: feature–target association and feature–feature redundancy. Pearson captures linear association; Spearman captures monotonic rank association. Neither establishes causality or guarantees predictive usefulness.
Experiment deliberately
Add the influential outlier, compare Pearson with Spearman, then run the feature selector step by step and explain why each candidate is kept or rejected.
Correlation measures association, not causation. Pearson correlation measures linear association; Spearman correlation measures monotonic association based on ranks. A feature can be important despite weak correlation when its relationship is nonlinear or interaction-dependent.
Correlation matrix
Inspect feature–feature redundancy as well as feature–target association.
Scatter + fitted line
Select a feature to inspect the raw relationship behind the coefficient.
Correlation-based feature selection
First require enough feature–target association, then reject features that are too correlated with an already selected stronger feature.
Ready.
PearsonUse for linear association between numeric variables; sensitive to influential outliers.
SpearmanUses ranks, so it captures monotonic relationships and is less driven by exact numeric spacing.
Feature selectionFit thresholds using training data only. Never use the final test set to choose features.
LimitationCorrelation screening is univariate: nonlinear effects, interactions and conditional usefulness can be missed.