Clustering, explained simply
So far every example came with a label. But sometimes nobody has labelled the data. Clustering is a type of machine learning that looks for groups on its own, by putting similar examples together.
A clustering model measures how close examples are to each other using their features, then forms groups of neighbours. You usually tell it how many groups to look for.
The model does not know what the groups mean. A person still has to look at them and decide whether they are useful: "these are all night-time birds" or "this group makes no sense".
Example: Sorting stars without names
- You have 30 stars with only a brightness and a colour.
- Ask for 3 clusters. The model groups bright blue stars, dim red stars and middling yellow stars.
- It never knew those names. You looked at the groups and gave them meaning.
Try this
- Tip out a box of buttons, beads or LEGO bricks and sort them into groups without being told how. Compare with someone else’s groups.
- Try asking for 2 groups, then 5 groups. Which feels more useful?
- Think of where a shop might use clustering, such as grouping customers who buy similar things.
