Labeling the data
As explained in the previous section, we have to manually label a training dataset. We assume that the sentiment is specific to the topic of analysis and that a person who labels the data is able to do it correctly. In order to store the labels, we create a column label, which associates a class with a tweet.
The general rule is that the more labeled data the better, but labeling is a costly operation. In order to be efficient we label only observations that are clearly associated with a class. It will help us to get only the best examples of positive, neutral, and negative sentiment. As a result, the algorithm should have a good performance in predicting the most insightful tweets in terms of sentiment, which is the objective. ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access