Follow the transformation
Choose features and scaling that define meaningful similarity.
Fit the clustering algorithm without target labels.
Inspect cluster size, separation/stability and representative observations.
K Means is an unsupervised clustering technique that groups observations using a specific notion of similarity or density.
K Means is an unsupervised clustering technique that groups observations using a specific notion of similarity or density. Clusters are model-dependent structures, not automatically “true” categories in the world.
K Means matters because unsupervised methods impose a particular notion of structure—variance, distance, density or probability—without target labels. The discovered representation or clusters must therefore be interpreted relative to that chosen notion.
Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Choose features and scaling that define meaningful similarity. Stage 2: Fit the clustering algorithm without target labels. Stage 3: Inspect cluster size, separation/stability and representative observations. Final checkpoint: Evaluate usefulness against the downstream analytical purpose.
Choose features and scaling that define meaningful similarity.
Fit the clustering algorithm without target labels.
Inspect cluster size, separation/stability and representative observations.
# Step 1 — Import the module so its functions/classes are available to the rest of this example.
import numpy as np
# Step 2 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.cluster import KMeans
# Step 3 — Construct `X` as an array so vectorised numerical operations can be applied consistently.
X = np.array([[1,1],[1.2,0.9],[5,5],[5.2,4.8]])
# Step 4 — Fit the model or transformer, learning its parameters from the supplied training data.
m = KMeans(n_clusters=2, n_init=10, random_state=0).fit(X)
# Step 5 — Display the current value explicitly so the result/state can be inspected during execution.
print(m.labels_)
# Step 6 — Display the current value explicitly so the result/state can be inspected during execution.
print(np.round(m.cluster_centers_, 2))Two label groups corresponding to the two compact clouds; centroids near [1.1, 0.95] and [5.1, 4.9].
For K Means, connect the displayed result to the specific input and mechanism above; independently verify one value/state change rather than treating successful execution as proof.
K-meansPartitions into k clusters by minimising within-cluster squared distance to centroids; favours roughly spherical equal-scale clusters.DBSCANDensity-based; finds arbitrary shapes and labels low-density points as noise; requires eps/min_samples.HierarchicalBuilds a nested dendrogram of merges/splits using a distance/linkage rule.Gaussian mixtureProbabilistic soft clustering with Gaussian component distributions.Use K Means when the goal is to discover or represent structure without a supervised target and the chosen similarity/density/latent assumptions match the data.
Reconsider the method when feature scaling, distance choice, cluster shape, density variation or embedding purpose makes the discovered structure unstable or uninterpretable.
Build a tiny, inspectable example of K Means. First choose features and scaling that define meaningful similarity. Then fit the clustering algorithm without target labels. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.
Before trusting a result from K Means, which check provides the strongest evidence that you understand and applied it correctly?
Step 1Choose features and scaling that define meaningful similarity.Step 2Fit the clustering algorithm without target labels.Step 3Inspect cluster size, separation/stability and representative observations.Step 4Visualise in original or reduced feature space without overclaiming separation.