Data Preparation for ML · Lesson 18

Feature Selection

Feature Selection changes or reduces the feature representation used by a model.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What Feature Selection actually means

Feature Selection changes or reduces the feature representation used by a model. Good feature engineering exposes relevant structure without leaking target or future information, while feature selection/dimensionality reduction control redundancy, noise and complexity.

Feature Selection matters because models learn from the feature representation they receive, not from the raw concept in your head. Scaling, encoding, imputation and construction can change geometry and signal, and learned steps must stay inside validation folds.

Deeper walkthrough

Read Feature Selection as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Start from a domain or modelling hypothesis about what information matters. Stage 2: Create the feature using only information available at the prediction time. Stage 3: Fit any learned feature transformation within the training fold. Final checkpoint: Keep a baseline to ensure added complexity actually helps.

Mechanism

Follow the transformation

Start from a domain or modelling hypothesis about what information matters.

Create the feature using only information available at the prediction time.

Fit any learned feature transformation within the training fold.

Evidence

Know what would convince you

  • Fit the transformation on training data only and confirm validation/test rows are transformed without refitting.
  • Inspect transformed shape, feature names/ranges and several rows against the raw inputs.
Useful distinctionFeature construction: Create new variables such as ratios, interactions or lags.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Start from a domain or modelling…

Start from a domain or modelling hypothesis about what information matters. For Feature Selection, identify the exact state before this stage, the operation or rule applied here, and the observable state afterwards so the mechanism remains inspectable.

State focus: identify exactly what changed at this stage and what observable evidence confirms that change.
How it works

Trace the mechanism step by step

  1. Start from a domain or modelling hypothesis about what information matters.
  2. Create the feature using only information available at the prediction time.
  3. Fit any learned feature transformation within the training fold.
  4. Evaluate whether the new representation improves validation performance or interpretability.
  5. Keep a baseline to ensure added complexity actually helps.
Worked demonstration

Make the concept concrete

Demonstration

Python / pandas example

# Step 1 — Import the module so its functions/classes are available to the rest of this example.
import pandas as pd
# Step 2 — Construct `df` as a tabular object with named columns for inspectable analysis.
df=pd.DataFrame({"x1":[1,2,3,4],"x2":[2,4,6,8],"noise":[4,1,3,2],"y":[1,2,3,4]})
# Step 3 — Display the current value explicitly so the result/state can be inspected during execution.
print(df.corr(numeric_only=True)["y"].sort_values(ascending=False))
Expected / illustrative result
x1 and x2 are both strongly related to y and also redundant with each other; selection should consider predictive evidence and redundancy inside validation.
Interpret the result.

For Feature Selection, trace representative source rows/columns into the result and reconcile row counts, dtypes, keys or missing values that the operation could change.

Distinctions & related ideas

Know what this is — and what it is not

Feature constructionCreate new variables such as ratios, interactions or lags.
Feature selectionKeep a subset of original/constructed variables.
PCACreate orthogonal linear combinations ordered by explained variance.
TF-IDFRepresent text by term importance relative to document frequency.
Use deliberately

When it is appropriate

Use Feature Selection when the model/analysis requires a deliberate representation of raw features and the transformation can be fit without leaking future or held-out information.

Boundary conditions

When to stop or reconsider

Avoid transformations that are unnecessary for the chosen model, cannot be reproduced at inference time, or learn from data that should remain held out.

Common mistakes

Failure modes to recognise

  • Fitting scaling, encoding, imputation or feature selection on the full dataset before validation.
  • Creating feature values whose meaning changes between training and inference.
  • Ignoring output shape/feature names and losing track of what the transformed columns represent.
Verification

How to check the result

  • Fit the transformation on training data only and confirm validation/test rows are transformed without refitting.
  • Inspect transformed shape, feature names/ranges and several rows against the raw inputs.
  • Place the transformation in a pipeline and cross-validate the complete workflow, not a preprocessed snapshot.
Feature importance & selection workflow

Calculate feature evidence before choosing what to keep

No single feature-importance number answers every question. Use statistical, univariate predictive and model-based evidence together, then validate the selected subset.

Inspect distributions→Welch t-test + p-value→Univariate ROC-AUC→Fit model→Permutation importance→RFE→Validate subset

What each method asks

  • Welch t-test / p-value: is the class-mean difference inconsistent with the equal-means null under the test assumptions?
  • AUC-based feature importance: use univariate max(AUC, 1-AUC) to measure one-feature discrimination, and/or measure the drop in model ROC-AUC when the fitted model loses that feature’s information.
  • Permutation importance: how much does a fitted model’s held-out performance fall when this feature is disrupted?
  • RFE: which features survive repeated model-based elimination, with subset size chosen using validation evidence?

Do not collapse the meanings

Important
A small p-value is not the same as a large predictive effect. A high univariate AUC does not prove the feature is needed in a multivariate model. RFE is model-dependent. None of these methods establishes causality.
Validation rule
Fit screening thresholds, scaling and RFE inside the training/validation workflow. Keep the final test set outside feature selection.
# Feature evidence on training data only
# Step 1 — Import the module so its functions/classes are available to the rest of this example.
import numpy as np
# Step 2 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from scipy.stats import ttest_ind
# Step 3 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.metrics import roc_auc_score
# Step 4 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.linear_model import LogisticRegression
# Step 5 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.feature_selection import RFE
# Step 6 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.inspection import permutation_importance

# Step 7 — Compute the right-hand expression and store its result in `feature_rows` for the next step.
feature_rows = []
# Step 8 — Iterate through the collection so the indented block is applied once for each item.
for name in feature_names:
    # Step 9 — Execute this statement and inspect how it changes the current value, object or program state.
    x0 = X_train.loc[y_train == 0, name].dropna()
    # Step 10 — Execute this statement and inspect how it changes the current value, object or program state.
    x1 = X_train.loc[y_train == 1, name].dropna()

    # 1) Statistical significance: unequal-variance two-sample test
    # Step 11 — Compute the right-hand expression and store its result in `t_stat, p_value` for the next step.
    t_stat, p_value = ttest_ind(x0, x1, equal_var=False)

    # 2) Univariate discrimination; direction-adjusted for importance
    # Step 12 — Compute the right-hand expression and store its result in `auc` for the next step.
    auc = roc_auc_score(y_train, X_train[name])
    # Step 13 — Compute the right-hand expression and store its result in `auc_importance` for the next step.
    auc_importance = max(auc, 1 - auc)

    # Step 14 — Execute this statement and inspect how it changes the current value, object or program state.
    feature_rows.append((name, t_stat, p_value, auc_importance))

# 3) Model-based subset selection: fit only on development data
# Step 15 — Instantiate `base` with the chosen algorithm/configuration before fitting it to data.
base = LogisticRegression(max_iter=2000)
# Step 16 — Compute the right-hand expression and store its result in `rfe` for the next step.
rfe = RFE(base, n_features_to_select=3)
# Step 17 — Fit the model or transformer, learning its parameters from the supplied training data.
rfe.fit(X_train_scaled, y_train)
# Step 18 — Construct `selected` as an array so vectorised numerical operations can be applied consistently.
selected = np.array(feature_names)[rfe.support_]

# 4) Model reliance: evaluate permutation loss on validation data
# Step 19 — Fit the model or transformer, learning its parameters from the supplied training data.
model = LogisticRegression(max_iter=2000).fit(
    X_train_scaled[:, rfe.support_], y_train
)
# Step 20 — Compute the right-hand expression and store its result in `perm` for the next step.
perm = permutation_importance(
    model,
    X_valid_scaled[:, rfe.support_],
    y_valid,
    scoring="roc_auc",
    n_repeats=20,
    random_state=42,
)

# Step 21 — Display the current value explicitly so the result/state can be inspected during execution.
print("Selected by RFE:", selected.tolist())
# Step 22 — Display the current value explicitly so the result/state can be inspected during execution.
print("Validation permutation ΔAUC:", perm.importances_mean.round(3))
Multiple features mean multiple tests

If many p-values are screened at once, consider false-discovery or family-wise error control and report effect sizes as well as significance. Use the p-value as evidence under assumptions, not as a mechanical feature-selection cutoff.

Correlation-based feature selection

Screen target signal, then remove redundant features

Correlation filtering is a fast filter method for numeric features. Calculate feature–target correlation to identify simple univariate signal, then inspect feature–feature correlation to avoid keeping multiple variables that carry almost the same information.

Training data only→Feature ↔ target correlation→Rank by |r|→Apply minimum signal threshold→Remove highly correlated duplicates→Validate selected subset
Useful forFast numeric screening and obvious redundancy reduction.
Do not assumeLow correlation means a feature is useless; nonlinear and interaction effects can be missed.
Leakage safeguardEstimate thresholds and feature choices inside the training/validation workflow—not on the final test set.
Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of Feature Selection. First start from a domain or modelling hypothesis about what information matters. Then create the feature using only information available at the prediction time. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Use a tiny train/validation split. Fit only on train, print the transformed shape/values, and confirm validation transformation does not update learned state.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from Feature Selection, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Start from a domain or modelling hypothesis about what information matters.
Step 2Create the feature using only information available at the prediction time.
Step 3Fit any learned feature transformation within the training fold.
Step 4Evaluate whether the new representation improves validation performance or interpretability.
Lesson summary

What to remember

  • Feature Selection changes or reduces the feature representation used by a model. Good feature engineering exposes relevant structure without leaking target or future information, while feature selection/dimensionality reduction control redundancy, noise and complexity.
  • Start from a domain or modelling hypothesis about what information matters.
  • Fitting scaling, encoding, imputation or feature selection on the full dataset before validation.
  • Fit the transformation on training data only and confirm validation/test rows are transformed without refitting.