November 2018
Intermediate to advanced
322 pages
7h 54m
English
Once you have installed the requisite packages (can be found in a requirements.txt file with the code) to run this project and read the data, the next step is to preprocess the data:
def get_processed_tokens(text):'''Gets Token List from a Review'''filtered_text = re.sub(r'[^a-zA-Z0-9\s]', '', text) #Removing Punctuationsfiltered_text = filtered_text.split()filtered_text = [token.lower() for token in filtered_text]return filtered_text
For example, if we have an input This is a GREAT movie!!!!, our output should be this is a great movie.
Read now
Unlock full access