July 2017
Beginner to intermediate
312 pages
7h 27m
English
First of all, we will use text analytics techniques to identify what are the most popular phrases related to technologies in repositories from 2017. Our analysis will be focused on the most frequent bigrams.
We import a nltk.collocation module which implements n-gram search tools:
import nltk from nltk.collocations import *
Then, we convert the clean description column into a list of tokens:
list_documents = df['clean'].apply(lambda x: x.split()).tolist()
As we perform an analysis on documents, we will use the method from_documents instead of a default one from_words. The difference between these two methods lies in the input data format. The one used in our case takes as argument a list of tokens and searches for n-grams ...
Read now
Unlock full access