5 · Data Preprocessing & Feature Engineering

Feature Selection

Feature Selection groups the core ideas a learner needs at the 5 · data preprocessing & feature engineering stage. Work through the lessons in order when new to the area, or use them independently as a reference when implementing an analysis.

How to use this topic

Learn the mechanism one decision at a time

Work through the lessons in order if the topic is new. If you already know the basics, open the specific leaf lesson that matches the operation, diagnostic or failure mode you need.

1Definition→
2Mechanism→
3Example→
4Diagnostic→
5Decision
01
Variance filteringVariance filtering is a practical concept within Feature Selection. It helps turn the broader workflow stage “5 · Data Preprocessing & Feature Engineering” into an explicit analytical decision that can be explained, implemented and checked. The concept should be understood in terms of purpose, mechanism, assumptions, evidence and downstream consequences.
02
Univariate statistical testsUnivariate statistical tests is a practical concept within Feature Selection. It helps turn the broader workflow stage “5 · Data Preprocessing & Feature Engineering” into an explicit analytical decision that can be explained, implemented and checked. The concept should be understood in terms of purpose, mechanism, assumptions, evidence and downstream consequences.
03
Mutual informationMutual information is a practical concept within Feature Selection. It helps turn the broader workflow stage “5 · Data Preprocessing & Feature Engineering” into an explicit analytical decision that can be explained, implemented and checked. The concept should be understood in terms of purpose, mechanism, assumptions, evidence and downstream consequences.
04
Recursive feature eliminationRecursive feature elimination changes how raw variables are represented for analysis or modelling. The transformation should preserve the information needed by the task while making assumptions explicit and reproducible.
05
Sequential feature selectionSequential feature selection changes how raw variables are represented for analysis or modelling. The transformation should preserve the information needed by the task while making assumptions explicit and reproducible.
06
L1 and model-based selectionL1 and model-based selection represents a family or practice in model building. The central idea is to define what structure can be learned, how model quality is measured during fitting, and how generalisation is tested on observations not used to choose the model.
07
RFECV and nested selectionRFECV and nested selection is a practical concept within Feature Selection. It helps turn the broader workflow stage “5 · Data Preprocessing & Feature Engineering” into an explicit analytical decision that can be explained, implemented and checked. The concept should be understood in terms of purpose, mechanism, assumptions, evidence and downstream consequences.