Mastering Machine Learning with R - Second Edition
by Cory Lesmeister, Doug Ortiz, Vikram Dhillon, Miroslav Kopecky
Hierarchical clustering
To build a hierarchical cluster model in R, you can utilize the hclust() function in the base stats package. The two primary inputs needed for the function are a distance matrix and the clustering method. The distance matrix is easily done with the dist() function. For the distance, we will use Euclidean distance. A number of clustering methods are available, and the default for hclust() is the complete linkage. We will try this, but I also recommend Ward's linkage method. Ward's method tends to produce clusters with a similar number of observations.
The complete linkage method results in the distance between any two clusters that is the maximum distance between any one observation in a cluster and any one observation ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access