August 2018
Intermediate to advanced
522 pages
12h 45m
English
To define the most frequently used impurity measures, we need to consider the total number of target classes:

In a certain node, j, we can define the probability p(y = i|Node = j) where i is an index [0, P-1] associated with each class. In other words, according to a frequentist approach, this value is the ratio between the number of samples belonging to class i and assigned to node j, and the total number of samples belonging to the selected node:

Read now
Unlock full access