July 2017
Beginner to intermediate
486 pages
13h 49m
English
The lossy count algorithm is used to identify elements in a dataset whose frequency count exceeds a user-given threshold. This algorithm takes data streams as an input instead of the finite set of a dataset.
With lossy counting, you periodically remove very low-count elements from the frequency table. The most frequently accessed words would almost never have low counts anyway, and if they did, they wouldn't be likely to stay there for long.
Here, the frequency threshold is usually defined by the user. When we give a parameter of min_count = 4, we remove the words that appear in the dataset less than four times and we will not consider them.
Read now
Unlock full access