Clustering Validation Metrics
Clustering evaluation measures compactness, separation or agreement with known labels while recognising that no single score defines a scientifically meaningful cluster. The topic is split into focused lessons so definitions, implementation decisions, diagnostics and common failure modes can be learned separately.
Learn the mechanism one decision at a time
Work through the lessons in order if the topic is new. If you already know the basics, open the specific leaf lesson that matches the operation, diagnostic or failure mode you need.
1Definition→
2Mechanism→
3Example→
4Diagnostic→
5Decision
Silhouette coefficientCompares within-cluster cohesion with separation from the nearest alternative cluster. Higher is generally better, but convex distance-based assumptions matter. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
02Davies–Bouldin indexMeasures average similarity between each cluster and its most similar alternative. Lower values indicate more separated compact clusters. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
03Calinski–Harabasz indexCompares between-cluster dispersion with within-cluster dispersion and often favours well-separated spherical structure. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
04External validationAdjusted Rand Index and Normalised Mutual Information compare clusters with known labels while correcting or normalising agreement. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
05StabilityRepeat clustering under resampling, initialisation or perturbation. Stable structure often matters more than marginal changes in one internal index. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.