July 2017
Beginner to intermediate
312 pages
7h 27m
English
Firstly, we use the code presented in previous chapters to find the most relevant bigrams in our dataset.
We import all the necessary libraries:
import nltk from nltk.collocations import * from nltk.corpus import stopwords import re
We define a function which will perform data cleaning and pre-processing tasks:
Finally, it will return a list of tokens:
def preprocess(text): #1)Basic cleaning text = text.strip() text = re.sub(r'https?:\/\/.*[\r\n]*', '',text, flags=re.MULTILINE) text = re.sub(r'[^\w\s]',' ',text) text = text.lower() #2) Tokenize single comment: ...
Read now
Unlock full access