Clustering – K-means

K-means is an unsupervised algorithm that creates K disjoint clusters of points with equal variance, minimizing the distortion (also named inertia).

Given only one parameter K, representing the number of clusters to be created, the K-means algorithm creates K sets of points S₁, S₂, …, S_K, each of them represented by its centroid: C₁, C₂, …, C_K. The generic centroid, C_i, is simply the mean of the samples of the points associated to the cluster Si in order to minimize the intra-cluster distance. The outputs of the system are as follows:

The composition of the clusters S₁, S₂, …, S_K, that is, the set of points composing the training set that are associated to the cluster number 1, 2, …, K.
The centroids of each cluster, C₁, C₂

Get Python: Real World Machine Learning now with the O’Reilly learning platform.

O’Reilly members experience books, live events, courses curated by job role, and more from O’Reilly and nearly 200 top publishers.

Start your free trial

Python: Real World Machine Learning by Prateek Joshi, John Hearty, Bastiaan Sjardin, Luca Massaron, Alberto Boschetti

Clustering – K-means

Don’t leave empty-handed

It’s yours, free.

Check it out now on O’Reilly