January 2018
Intermediate to advanced
310 pages
7h 48m
English
Vinyals et al., in the paper https://arxiv.org/pdf/1411.4555.pdf, proposed an end to end trainable deep learning for image captioning, which has CNN and RNN stacked back to back. This is an end to end trainable model. The structure is shown here:

This model could generate a sentence that is completed in natural language. The expanded view of the CNN and LSTM is shown here:

Read now
Unlock full access