k-means: watch a machine find groups with nobody to teach it

Scatter some points, pick how many groups to look for, and watch centroids drift until the clusters snap into place. No labels, no teacher: this is unsupervised learning you can watch.

y26w37

Last week the machine rolled downhill toward a known answer. This week nobody tells it the answer at all. You hand it a pile of points and a single number, k, the number of groups to look for. It has to discover the groups itself. This is unsupervised learning: no labels, no teacher, just structure the machine digs out of raw data.

The gadget below is a k-means playground. The small dots are your data. The big diamonds are centroids, the current guess for the middle of each group. The whole algorithm is two moves repeated:

  1. Assign: every point takes the color of its nearest centroid.
  2. Update: every centroid slides to the average position of the points that just chose it.

That is it. Do those two moves over and over (this is called Lloyd’s algorithm) and the centroids drift until nobody wants to switch groups anymore.

Try this:

  1. Press Step a few times. Watch the two phases alternate: points recolor, then the diamonds jump to the middle of their color. Press Run to loop until it settles.
  2. Click empty space to add your own points, or press New points for a fresh scatter. Then run it again and watch new borders form.
  3. Change k. Ask for 2 groups where there are clearly 3, and it is forced to merge two blobs. Ask for 5 where there are 3, and it splits a real group in half. k-means finds exactly as many groups as you ask for, whether or not they are really there.

Why this matters

Most of the data in the world arrives with no labels: web pages, images, customer behavior, gene expression. You cannot always tell a machine the right answer, because nobody knows it. Clustering is how a system carves that raw pile into structure on its own, and k-means, published in the 1950s, is still the first tool most people reach for.

It is the counterpart to supervised learning, where the model trains on labeled examples. Both are branches of machine learning, but unsupervised learning works without an answer key. The same instinct, “group similar things together,” runs underneath the giant datasets that trained modern vision and language models: before a network could learn to label the ImageNet photos that powered AlexNet, someone had to organize a mountain of unlabeled images. Behind the magic is a simple loop: color the points, move the middles, repeat.

Missed the earlier editions? Train a perceptron, watch a Markov chain babble, lose to unbeatable tic-tac-toe, teach a filter to see edges, watch an agent learn from reward, see how a model reads text in tokens, or watch one roll downhill to learn.

Groups (k)

Click the box to add points, pick how many groups k to look for, then press Step. Each step first recolors every point to its nearest centroid (the big markers), then moves each centroid to the middle of its group. Press Run to loop until it settles. No labels, no teacher: this is unsupervised learning.