Data Preparation for ML · Lesson 19

PCA

Principal Component Analysis (PCA) is an unsupervised linear transformation that rotates centred data into orthogonal directions of decreasing variance.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What PCA actually means

Principal Component Analysis (PCA) is an unsupervised linear transformation that rotates centred data into orthogonal directions of decreasing variance. Each principal component is a weighted linear combination of the original features; projecting onto the first components can compress correlated data while preserving as much variance as possible under a linear criterion.

PCA matters because models learn from the feature representation they receive, not from the raw concept in your head. Scaling, encoding, imputation and construction can change geometry and signal, and learned steps must stay inside validation folds.

Deeper walkthrough

Read PCA as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Standardise features first when different measurement scales should contribute comparably. Stage 2: Centre the training data. Stage 3: Compute covariance structure (or use SVD) and obtain orthogonal component directions. Final checkpoint: Interpret loadings and reconstruction/variance trade-offs; fit PCA inside cross-validation.

Mechanism

Follow the transformation

Standardise features first when different measurement scales should contribute comparably.

Centre the training data.

Compute covariance structure (or use SVD) and obtain orthogonal component directions.

Evidence

Know what would convince you

  • Fit the transformation on training data only and confirm validation/test rows are transformed without refitting.
  • Inspect transformed shape, feature names/ranges and several rows against the raw inputs.
Useful distinctionFeature construction: Create new variables such as ratios, interactions or lags.
Visual demonstration of PCA
Visual demonstration: use the diagram to trace the main objects and state changes involved in PCA.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Standardise features first when different measurement…

Standardise features first when different measurement scales should contribute comparably. For PCA, identify the exact state before this stage, the operation or rule applied here, and the observable state afterwards so the mechanism remains inspectable.

State focus: identify exactly what changed at this stage and what observable evidence confirms that change.
Mathematical / formal view
For centred matrix X, PCA finds orthonormal directions w that maximise Var(Xw), equivalently the eigenvectors of the covariance matrix / right singular vectors of X.
How it works

Trace the mechanism step by step

  1. Standardise features first when different measurement scales should contribute comparably.
  2. Centre the training data.
  3. Compute covariance structure (or use SVD) and obtain orthogonal component directions.
  4. Order components by explained variance.
  5. Project observations onto the chosen components using the training-fitted transformation.
  6. Interpret loadings and reconstruction/variance trade-offs; fit PCA inside cross-validation.
Worked demonstration

Make the concept concrete

Demonstration

Python / scikit-learn example

# Step 1 — Import the module so its functions/classes are available to the rest of this example.
import numpy as np
# Step 2 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.decomposition import PCA
# Step 3 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.preprocessing import StandardScaler
# Step 4 — Construct `X` as an array so vectorised numerical operations can be applied consistently.
X = np.array([[1,10],[2,20],[3,31],[4,39],[5,51]], dtype=float)
# Step 5 — Fit the transformation on the training input and immediately transform that same input.
Xs = StandardScaler().fit_transform(X)
# Step 6 — Fit the model or transformer, learning its parameters from the supplied training data.
pca = PCA(n_components=1).fit(Xs)
# Step 7 — Apply the already-fitted transformation without relearning its parameters from this data.
Z = pca.transform(Xs)
# Step 8 — Display the current value explicitly so the result/state can be inspected during execution.
print("explained variance ratio:", np.round(pca.explained_variance_ratio_, 3))
# Step 9 — Display the current value explicitly so the result/state can be inspected during execution.
print("scores shape:", Z.shape)
Expected / illustrative result
explained variance ratio: [about 0.999]
scores shape: (5, 1)
Interpret the result.

For PCA, trace representative source rows/columns into the result and reconcile row counts, dtypes, keys or missing values that the operation could change.

Distinctions & related ideas

Know what this is — and what it is not

Feature constructionCreate new variables such as ratios, interactions or lags.
Feature selectionKeep a subset of original/constructed variables.
PCACreate orthogonal linear combinations ordered by explained variance.
TF-IDFRepresent text by term importance relative to document frequency.
Use deliberately

When it is appropriate

Use PCA when the model/analysis requires a deliberate representation of raw features and the transformation can be fit without leaking future or held-out information.

Boundary conditions

When to stop or reconsider

Avoid transformations that are unnecessary for the chosen model, cannot be reproduced at inference time, or learn from data that should remain held out.

Common mistakes

Failure modes to recognise

  • Fitting scaling, encoding, imputation or feature selection on the full dataset before validation.
  • Creating feature values whose meaning changes between training and inference.
  • Ignoring output shape/feature names and losing track of what the transformed columns represent.
Verification

How to check the result

  • Fit the transformation on training data only and confirm validation/test rows are transformed without refitting.
  • Inspect transformed shape, feature names/ranges and several rows against the raw inputs.
  • Place the transformation in a pipeline and cross-validate the complete workflow, not a preprocessed snapshot.
Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of PCA. First standardise features first when different measurement scales should contribute comparably. Then centre the training data. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Use a tiny train/validation split. Fit only on train, print the transformed shape/values, and confirm validation transformation does not update learned state.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from PCA, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Standardise features first when different measurement scales should contribute comparably.
Step 2Centre the training data.
Step 3Compute covariance structure (or use SVD) and obtain orthogonal component directions.
Step 4Order components by explained variance.
Lesson summary

What to remember

  • Principal Component Analysis (PCA) is an unsupervised linear transformation that rotates centred data into orthogonal directions of decreasing variance. Each principal component is a weighted linear combination of the original features; projecting onto the first components can compress correlated data while preserving as much variance as possible under a linear criterion.
  • Standardise features first when different measurement scales should contribute comparably.
  • Fitting scaling, encoding, imputation or feature selection on the full dataset before validation.
  • Fit the transformation on training data only and confirm validation/test rows are transformed without refitting.