What about cross-validation?
Often, in the case of smaller datasets, data scientists employ a technique known as cross-validation, which is also available to you in Spark. The CrossValidator class starts by splitting the dataset into N-folds (user declared) - each fold is used N-1 times as part of the training set and once for model validation. For example, if we declare that we wish to use a 5-fold cross-validation, the CrossValidator class will create five pairs (training and testing) of datasets using four-fifths of the dataset to create the training set with the final fifth as the test set, as shown in the following figure.
The idea is that we would see the performance of our algorithm across different, randomly sampled datasets to account ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access