July 2019
Intermediate to advanced
512 pages
19h 39m
English
Define a function for preprocessing the dataset:
def pre_process(text): # convert to lowercase text = str(text).lower() # remove all special characters and keep only alpha numeric characters and spaces text = re.sub(r'[^A-Za-z0-9\s.]',r'',text) #remove new lines text = re.sub(r'\n',r' ',text) # remove stop words text = " ".join([word for word in text.split() if word not in stopWords]) return text
We can see how the preprocessed text looks like by running the following code:
pre_process(data[0][50])
We get the output as:
'agree fancy. everything needed. breakfast pool hot tub nice shuttle airport later checkout time. noise issue tough sleep through. awhile forget noisy door nearby noisy guests. ...
Read now
Unlock full access