July 2017
Beginner to intermediate
312 pages
7h 27m
English
For the first kind, we have to create a new variable which contains a cleaned string. We will do it in three steps which have already been presented in previous chapters:
As we work only on English data, we should remove all the descriptions which are written in other languages. The main reason to do so is that each language requires a different processing and analysis flow. If we left descriptions in Russian or Chinese, we would have very noisy data which we would not be able to interpret. As a consequence, we can say that we are analyzing trends in the English-speaking world.
Firstly, we remove all the empty strings in the description column.
df = df.dropna(subset=['description']) ...
Read now
Unlock full access