Chapter 13Batchin’ Up
By now, we’re familiar enough with gradient descent. This chapter introduces a souped-up variant of GD: mini-batch gradient descent.
Mini-batch gradient descent is slightly more complicated than plain vanilla GD—but as we’re about to see, it also tends to converge faster. In simpler terms, mini-batch GD is faster at approaching the minimum loss, speeding up the network’s training. As a bonus, it takes less memory, and sometimes it even finds a better loss than regular GD. In fact, after this chapter, you might never use regular GD again!
You might wonder why we’re focusing on training speed, when we have more pressing concerns to deal with. In particular, the accuracy of our neural network is still disappointing—better ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access