August 2018
Intermediate to advanced
522 pages
12h 45m
English
An unnormalized dataset with many features contains information proportional to the independence of all features and their variance. Let's consider a small dataset with three features, generated with random Gaussian distributions:

Even without further analysis, it's obvious that the central line (with the lowest variance) is almost constant and doesn't provide any useful information. Recall from Chapter 2, Important Elements in Machine Learning, that the entropy H(X) is quite small, while the other two variables carry more information. ...
Read now
Unlock full access