July 2019
Intermediate to advanced
512 pages
19h 39m
English
Now, we will see how to perform document classification using doc2vec. In this section, we will use the 20 news_dataset. It consists of 20,000 documents over 20 different news categories. We will use only four categories: Electronics, Politics, Science, and Sports. So, we have 1,000 documents under each of these four categories. We rename the documents with a prefix, category_. For example, all science documents are renamed as Science_1, Science_2, and so on. After renaming them, we combine all the documents and place them in a single folder. The combined data, along with complete code is available at available as a Jupyter Notebook on GitHub at http://bit.ly/2KgBWYv.
Now, we train our doc2vec model ...
Read now
Unlock full access