0 / 9
K-Means Clustering
A way to split data with no right answer (no group labels) into K groups. Each group has one center. Every data point joins the group of its nearest center, then each center moves to the average position of the points in its group. This repeats until the centers stop moving. K is the number of groups, and "means" refers to placing each center at the average position.
Example: plot 12 café customers by number of visits and spend per visit, then split similar customers into three groups.
1 / 9
Start › Choose K
Set the number of groups (clusters) K to 3. The algorithm does not pick K for you, so the analyst decides it first.
2 / 9
Start › Pick starting centers
Pick 3 data points at random as the first centers (center 1, 2 and 3). This time one came from each group.
3 / 9
Process › Measure distances
Measure the straight-line distance from every customer to each of the three centers. The diagram shows one customer at the top right: 6.6 to center 1, 1.2 to center 2 and 5.9 to center 3.
4 / 9
Process › Assign to clusters
Put each customer in the group of the nearest center. The example customer is closest to center 2, so it joins center 2's group. The 12 customers now form three groups.
5 / 9
Process › Recompute centers
For each group, find the average position of its customers and move the center there. The starting centers sat on a customer, but the moved centers land in the empty middle of each group.
6 / 9
End › Compare centers
Compare the moved centers (new) with the centers before the move (old, dashed squares). They changed position, so instead of stopping, go back to the process step and run it again.
7 / 9
Process › Distances and assignment (pass 2)
Measure distances again from the new centers and reassign each customer to the nearest one. This time no customer moves to another group.
8 / 9
Process › Recompute centers (pass 2)
The groups are unchanged, so their average positions are unchanged too. The centers do not move.
9 / 9
End › Compare centers
The new centers are the same as the old ones. Nothing else will change, so the algorithm stops here with three groups: top left, top right and bottom right.
1 / 9
Start › Choose K
Set the number of groups (clusters) K to 3, the same as the first tab.
2 / 9
Start › Pick starting centers
Three data points were picked at random: two (center 1 and 2) from the top-left group and one (center 3) from the bottom-right group. The top-right group has no center.
3 / 9
Process › Measure distances
For one customer at the top right, the distances are 6.7 to center 1, 5.8 to center 2 and 5.2 to center 3. No center is close, so center 3 is the nearest of the three.
4 / 9
Process › Assign to clusters
The top-left group splits between center 1 and center 2, and all 8 customers at the top right and bottom right join center 3's group.
5 / 9
Process › Recompute centers
Center 3 moves to the average position of those 8 customers, landing in the empty space between the two groups. Centers 1 and 2 shift a little inside the top-left group.
6 / 9
End › Compare centers
The centers changed position, so instead of stopping, go back to the process step and run it again.
7 / 9
Process › Distances and assignment (pass 2)
Measuring the distances again moves no customer to another group. For the customers at the top right, center 3 is still the nearest.
8 / 9
Process › Recompute centers (pass 2)
The groups are unchanged, so the centers do not move.
9 / 9
End › Compare centers
The new centers match the old ones, so it stops. But the top-left group was split in two, and the two groups on the right were merged into one. K-means can stop at a different result depending on where the starting centers are picked (a local optimum). That is why it is run several times with different starting centers and the best grouping is kept.