August 2018
Intermediate to advanced
522 pages
12h 45m
English
The most useful measure in information theory (as well as in machine learning) is called entropy:

This value is proportional to the uncertainty of X and is measured in bits (if the logarithm has another base, this unit can change too). For many purposes, a high entropy is preferable, because it means that a certain feature contains more information. For example, in tossing a coin (two possible outcomes), H(X) = 1 bit, but if the number of outcomes grows, even with the same probability, H(X) also does because of a higher number of different values, and therefore has increased variability. It's possible to prove that for a Gaussian distribution ...
Read now
Unlock full access