July 2019
Intermediate to advanced
512 pages
19h 39m
English
Let's calculate the gradients of loss with respect to hidden-to-input layer weights
for all the gates and the candidate state. Computing gradients of loss with respect to
is exactly the same as the gradients we computed with respect to
, except that the last term will be
instead of . Let's examine what we mean by that. ...
Read now
Unlock full access