July 2019
Intermediate to advanced
512 pages
19h 39m
English
After predicting the output,
, we are in the final layer of the network. Since we are backpropagating; that is, going from the output layer to the input layer, our first weight will be
, which is hidden-to-output layer weight.
We have learned throughout that the final loss is the sum of the loss over all the time steps. In a similar manner, our final gradient is the sum of gradients at all time steps as follows:

If ...
Read now
Unlock full access