book

40 Algorithms Every Programmer Should Know

by Imran Ahmad

June 2020

Intermediate to advanced

382 pages

11h 39m

English

Packt Publishing

Read now

Unlock full access

Content preview from 40 Algorithms Every Programmer Should Know

Tokenization

When we are working with NLP, the first job is to divide the text into a list of tokens. This process is called tokenization. The granularity of the resulting tokens will vary based on the objective—for example, each token can consist of the following:

A word
A combination of words
A sentence
A paragraph

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Start your free trial

50 Algorithms Every Programmer Should Know - Second Edition

Imran Ahmad

Grokking Algorithms

Aditya Bhargava

Ultimate Go Programming, Second Edition

William Kennedy

Data Structures and Algorithms: The Complete Masterclass

Shubham Sarda

Publisher Resources

ISBN: 9781789801217Supplemental Content