July 2017
Beginner to intermediate
486 pages
13h 49m
English
Doc2vec (Document vectors) is an extension of word2vec. It learns to correlate document labels and words, rather than words with other words. Here, you need document tags. You are able to represent an entire sentence using a fixed-length vector. This is also using word2vec concepts. If you feed the sentences with labels into the neural network, then it performs classification on a given dataset. So, in short, you tag your text and then use this tagged dataset as input and apply the Doc2vec technique on that given dataset. This algorithm will generate tag vectors for the given text. You can find the code at this GitHub link: https://github.com/jalajthanaki/NLPython/blob/master/ch6/doc2vecexample.py
Read now
Unlock full access