April 2025
Intermediate
384 pages
11h 15m
English
In this appendix, we present expressions for computing the partial derivatives of loss with respect to RNN parameters of
and
during the back propagation through time used in RNNs as discussed in Chapter 10. Let us first present the definitions of RNN parameters from Chapter 10:
where
is realized as tanh and
as softmax. In order to minimize the loss
, its gradients with respect to and must be computed. Recall from Eq. (10.7) in Chapter 10 that cross‐entropy loss is defined as ...
Read now
Unlock full access