Let's look at the following optimizers:
- Stochastic gradient descent (SGD): This is the simplest optimizer. For every weight, the model stores an error rate. Depending on the direction of the error, whether the predicted value is greater than or less than the true value, a small learning rate is applied to the weight to change the next round of results so that the predicted values are incrementally moving in the opposite direction as the error. This process is simple, but for deep neural networks, the fine adjustments at every iteration mean that it can take a long time for the model to converge.
- Momentum: Momentum itself is not an optimizer, but is something that can be incorporated with different optimizers. The idea of momentum ...