
CHAPTER 2
Similarity Join for Big
Geographic Data
Yasin N. Silva,*
Jason M. Reed,
Lisa M. Tsosie and
Timothy A. Matti
Introduction
Similarity Join is one of the most useful data processing and analysis
operations for geographic data. It retrieves all data pairs whose distances are
smaller than a predefi ned threshold ε. Multiple application scenarios need
to perform this operation over large amounts of data. Internet companies,
for instance, collect massive amounts of information on their customers such
as their geographic location and interests. They can use similarity queries
to provide enhanced services to their customers; for example, a movie ...