July 2017
Beginner to intermediate
486 pages
13h 49m
English
In raw text data, data is in paragraph form. Now, if you want the sentences from the paragraph, then you need to tokenize at sentence level.
Some of the specialized cases need a customized rule for the sentence tokenizer as well.
The following open source tools are available for performing sentence tokenization:
Here we are using the nltk sentence tokenizer.
We are using sent_tokenize from nltk and will import it as st:
Read now
Unlock full access