Discover structure with k-means
Group unlabeled points and learn why clusters need interpretation.
- Alternate assignment and centroid updates
- Handle an empty cluster
- Separate geometric structure from meaning
Group by a chosen geometry
K-means seeks k centers that reduce squared distances between points and their assigned center. It alternates between assigning each point to its nearest center and moving each center to the mean of its assigned points.
The method does not discover authoritative categories. The number of centers, initialization, scaling, and distance assumptions shape the result. Two clusters may simply reflect the units or collection process rather than a useful difference.
Inspect more than the centers
Some initializations leave a cluster empty. Keep its previous center or use a documented reinitialization strategy. Multiple initializations can find different local solutions; record how you selected a run.
Look at actual members of each cluster and compare sizes. K-means favors compact groups and struggles with curved shapes, unequal densities, and outliers. Unsupervised does not mean free of assumptions.
A small experiment you can run.
The centers summarize two obvious synthetic groups. The implementation preserves an old center if its assigned group is empty.
from math import dist
points = [(0., 0.), (0., 2.), (2., 0.), (8., 8.), (8., 10.), (10., 8.)]
centers = [points[0], points[-1]]
for _ in range(10):
groups = [[], []]
for point in points:
nearest = min(range(2), key=lambda i: dist(point, centers[i]))
groups[nearest].append(point)
centers = [tuple(sum(p[d] for p in group)/len(group) for d in range(2))
if group else centers[i] for i, group in enumerate(groups)]
print("Centers:", centers)
print("Group sizes:", [len(group) for group in groups])
Save the file, open your terminal in that folder, and run python clustering-with-k-means.py. Use python3 or py if required by your installation. Setup guide
The original groups contain three points each.
Observe an outlier.
- Add the point (100, 100).
- Run with the same initial centers.
- Compare the new centers and group sizes with the original result.
Compare with a suggested solution
The distant point can pull a center away from the ordinary group. Squared-distance objectives penalize long distances strongly. Investigate outliers before deleting them: they may be errors or important rare examples.
One idea to take with you.
Make it part of your progress.
Finish the practice and answer the knowledge check to mark this lesson complete.
Go deeper with primary documentation
Optional references for further study. This lesson and its examples were written for Artificials.
scikit-learn: model evaluationscikit-learn: common pitfalls