January 2018
Intermediate to advanced
310 pages
7h 48m
English
Mao et al., in the paper https://arxiv.org/pdf/1412.6632.pdf, proposed a method that uses multimodal embedding space to generate the captions. The following figure illustrates this approach:

Kiros et al., in the paper https://arxiv.org/pdf/1411.2539.pdf, proposed another multimodal approach to generate captions, which can embed both image and text into the same multimodal space. The following figure illustrates this approach:

Both of the multimodal approaches ...
Read now
Unlock full access