July 2019
Intermediate to advanced
422 pages
13h 57m
English
Raul Gomez⁎,†; Lluis Gomez†; Jaume Gibert⁎; Dimosthenis Karatzas† ⁎Eurecat, Centre Tecnològic de Catalunya, Unitat de Tecnologies Audiovisuals, Barcelona, Spain†Computer Vision Center, Universitat Autònoma de Barcelona, Barcelona, Spain
Self-supervised learning from multimodal image and text data allows deep neural networks to learn powerful features with no need of human-annotated data. Web and social media platforms provide a virtually unlimited amount of this multimodal data. In this work we propose to exploit this free available data to learn a multimodal image and text embedding, aiming to leverage the semantic knowledge learned in the text domain and transfer ...
Read now
Unlock full access