August 2019
Beginner
482 pages
12h 56m
English
Decision trees, introduced in Chapter 13, Training a Machine Learning Model, and which we have been using so far, are fast and easy to interpret. Their weak point, however, is overfitting—many features might seem to be a great predictor on the training dataset, but turn out to mislead the models on the external data. In other words, they don't represent the general population. The problem is that decision trees (another algorithm) don't have any internal mechanics to detect and ignore those features.
A suite of more sophisticated models was developed on top of decision models to fight overfitting. These models are usually called tree ensembles, as all of them train multiple decision trees and aggregate their predictions. ...