July 2019
Intermediate to advanced
512 pages
19h 39m
English
The total loss,
, is the sum of losses at all time steps, and can be given as follows:

To minimize the loss using gradient descent, we find the derivative of loss with respect to all of the weights used in the GRU cell as follows:
, which are the input-to-hidden layer weights of the update gate, reset gate, and content state, respectivelyRead now
Unlock full access