February 2019
Intermediate to advanced
386 pages
9h 54m
English
This measure (together with all the other ones discussed from now on) is based on knowledge of the ground truth. Before introducing the index, it's helpful to define some common values. If we denote with Ytrue the set containing the true assignments and with Ypred, the set of predictions (both containing M values and K clusters), we can estimate the following probabilities:

In the previous formulas, ntrue/pred(k) represents the number of true/predicted samples belonging the cluster k ∈ K. At this point, we can compute the entropies of Ytrue and Ypred:
Considering the definition of entropy, H(•) is maximized by a uniform ...
Read now
Unlock full access