book

Python Data Analysis Cookbook

by Ivan Idris

July 2016

Beginner to intermediate

462 pages

9h 14m

English

Packt Publishing

Read now

Unlock full access

Content preview from Python Data Analysis Cookbook

Tokenizing news articles in sentences and words

The corpora that are part of the NLTK distribution are already tokenized, so we can easily get lists of words and sentences. For our own corpora, we should apply tokenization too. This recipe demonstrates how to implement tokenization with NLTK. The text file we will use is in this book's code bundle. This particular text is in English, but NLTK supports other languages too.

Getting ready

Install NLTK, following the instructions in the Introduction section of this chapter.

How to do it...

The program is in the tokenizing.py file in this book's code bundle:

The imports are as follows:

from nltk.tokenize import sent_tokenize
from nltk.tokenize import word_tokenize
import dautil as dl

The following code demonstrates ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Start your free trial

Python Machine Learning Cookbook - Second Edition

Giuseppe Ciaburro, Prateek Joshi

Python: End-to-end Data Analysis

Phuong Vothihong, Martin Czygan, Ivan Idris, Magnus Vilhelm Persson, Luiz Felipe Martins

Practical Data Analysis Cookbook

Tomasz Drabas

Python Data Science Essentials - Third Edition

Alberto Boschetti, Luca Massaron, Pietro Marinelli, Matteo Malosetti

Publisher Resources

ISBN: 9781785282287Supplemental Content