Unsupervised Learning · Lesson 50

Hierarchical Clustering

Hierarchical Clustering is an unsupervised clustering technique that groups observations using a specific notion of similarity or density.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What Hierarchical Clustering actually means

Hierarchical Clustering is an unsupervised clustering technique that groups observations using a specific notion of similarity or density. Clusters are model-dependent structures, not automatically “true” categories in the world.

Hierarchical Clustering matters because unsupervised methods impose a particular notion of structure—variance, distance, density or probability—without target labels. The discovered representation or clusters must therefore be interpreted relative to that chosen notion.

Deeper walkthrough

Read Hierarchical Clustering as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Choose features and scaling that define meaningful similarity. Stage 2: Fit the clustering algorithm without target labels. Stage 3: Inspect cluster size, separation/stability and representative observations. Final checkpoint: Evaluate usefulness against the downstream analytical purpose.

Mechanism

Follow the transformation

Choose features and scaling that define meaningful similarity.

Fit the clustering algorithm without target labels.

Inspect cluster size, separation/stability and representative observations.

Evidence

Know what would convince you

  • Standardise/represent features deliberately and document the distance or latent objective.
  • Repeat with nearby hyperparameters/seeds or resamples and examine whether major structure is stable.
Useful distinctionK-means: Partitions into k clusters by minimising within-cluster squared distance to centroids; favours roughly spherical equal-scale clusters.
Visual demonstration of Hierarchical Clustering
Visual demonstration: use the diagram to trace the main objects and state changes involved in Hierarchical Clustering.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Choose features and scaling that define…

Choose features and scaling that define meaningful similarity. This is an input-preparation stage for Hierarchical Clustering. Verify the relevant type, shape, units, keys, missingness or assumptions before later steps depend on them.

Input focus: confirm the data/object, units, type, shape and assumptions before the next operation depends on them.
How it works

Trace the mechanism step by step

  1. Choose features and scaling that define meaningful similarity.
  2. Fit the clustering algorithm without target labels.
  3. Inspect cluster size, separation/stability and representative observations.
  4. Visualise in original or reduced feature space without overclaiming separation.
  5. Evaluate usefulness against the downstream analytical purpose.
Worked demonstration

Make the concept concrete

Demonstration

Text example

Start with each point as its own cluster.
Repeatedly merge the pair of clusters chosen by the linkage rule.
The dendrogram records merge order and height; cutting it at a chosen height yields a flat clustering.
Expected / illustrative result
Different linkage rules can produce different hierarchies because “distance between clusters” is defined differently.
Interpret the result.

For Hierarchical Clustering, connect the displayed result to the specific input and mechanism above; independently verify one value/state change rather than treating successful execution as proof.

Distinctions & related ideas

Know what this is — and what it is not

K-meansPartitions into k clusters by minimising within-cluster squared distance to centroids; favours roughly spherical equal-scale clusters.
DBSCANDensity-based; finds arbitrary shapes and labels low-density points as noise; requires eps/min_samples.
HierarchicalBuilds a nested dendrogram of merges/splits using a distance/linkage rule.
Gaussian mixtureProbabilistic soft clustering with Gaussian component distributions.
Use deliberately

When it is appropriate

Use Hierarchical Clustering when the goal is to discover or represent structure without a supervised target and the chosen similarity/density/latent assumptions match the data.

Boundary conditions

When to stop or reconsider

Reconsider the method when feature scaling, distance choice, cluster shape, density variation or embedding purpose makes the discovered structure unstable or uninterpretable.

Common mistakes

Failure modes to recognise

  • Treating discovered clusters/components/embeddings as objectively true categories.
  • Ignoring feature scale and distance/metric choices that dominate the geometry.
  • Selecting a visually pleasing result without checking stability or external/domain plausibility.
Verification

How to check the result

  • Standardise/represent features deliberately and document the distance or latent objective.
  • Repeat with nearby hyperparameters/seeds or resamples and examine whether major structure is stable.
  • Check cluster/component/embedding summaries against original variables and domain expectations rather than relying on the plot alone.
Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of Hierarchical Clustering. First choose features and scaling that define meaningful similarity. Then fit the clustering algorithm without target labels. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Use a small synthetic dataset where you know the rough structure. Change one scale or hyperparameter and observe which parts of the result are stable.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from Hierarchical Clustering, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Choose features and scaling that define meaningful similarity.
Step 2Fit the clustering algorithm without target labels.
Step 3Inspect cluster size, separation/stability and representative observations.
Step 4Visualise in original or reduced feature space without overclaiming separation.
Lesson summary

What to remember

  • Hierarchical Clustering is an unsupervised clustering technique that groups observations using a specific notion of similarity or density. Clusters are model-dependent structures, not automatically “true” categories in the world.
  • Choose features and scaling that define meaningful similarity.
  • Treating discovered clusters/components/embeddings as objectively true categories.
  • Standardise/represent features deliberately and document the distance or latent objective.