December 2018
Beginner to intermediate
616 pages
15h 56m
English
Let's assume that we are in a binary classification problem setting and want to use RandomForestClassifier. All SparkML algorithms have a compatible API, so they can be used interchangeably. So it really doesn't matter which one we use, but RandomForestClassifier has more (hyper)parameters than more simple models like logistic regression. At a later stage we'll use (hyper)parameter tuning which is also inbuilt in Apache SparkML. Therefore it makes sense to use an algorithm where more knobs can be tweaked. Adding such a binary classifier to our Pipeline is very simple:
import org.apache.spark.ml.classification.RandomForestClassifiervar rf = new RandomForestClassifier() .setLabelCol("label") .setFeaturesCol("features") ...Read now
Unlock full access