Skip to Content
算法技术手册(原书第2 版)
book

算法技术手册(原书第2 版)

by George T.Heineman, Gary Pollice, Stanley Selkow
August 2017
Intermediate to advanced
360 pages
8h 35m
Chinese
China Machine Press
Content preview from 算法技术手册(原书第2 版)
搜索算法
111
搜索算法
5.4 布隆过滤器
无论是使用链表还是开放定址技术,
散列搜索
都需要集合
C
中的所有元素存储到散列
H
中,而且随着更多的元素加入散列表中,如果不增加存储空间,那么势必会需要花
费更多的时间来寻找元素。话虽如此,本章所有算法的性能也都和存储空间的大小息息
相关,我们希望能够寻找到一个平衡点——在从集合中查找元素时,既能够减少查找的
次数,也能够减少额外需要的空间。
布隆过滤器
提供了另外一种思路,它使用一个
位数组
B
来确保只需要
常数时间
就能够
把集合
C
中的元素插入集合
B
中,或者检查一个元素是否没有被添加到
B
。而且令人
惊讶的是,这个常数时间和
B
中已有的元素数量无关。当然世事并非完美,
布隆过滤
在检查一个元素是否在集合
B
中时,
误报
目标元素存在,然而实际并不存在。相反,
布隆过滤器
却可以很准确地汇报目标元素并不存在
B
中。
布隆过滤器算法小结
最好情况、平均情况和最坏情况:
O
(
k
)
create(m)
return bit array of m bits
end
add (bits,value)
foreach hashFunction hf
setbit = 1 << hf(value)
bits |= setbit
end
search (bits,value)
foreach hashFunction hf
checkbit = 1 << hf(value)
if checkbit | bits = 0 then
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

AirBnbBlueOriginElectronic ArtsHomeDepotNasdaqRakutenTata Consultancy Services

QuotationMarkO’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.
Julian F.
Head of Cybersecurity
QuotationMarkI wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.
Addison B.
Field Engineer
QuotationMarkI’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.
Amir M.
Data Platform Tech Lead
QuotationMarkI'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.
Mark W.
Embedded Software Engineer

You might also like

机器学习实战:基于Scikit-Learn、Keras 和TensorFlow (原书第2 版)

机器学习实战:基于Scikit-Learn、Keras 和TensorFlow (原书第2 版)

Aurélien Géron
Go语言编程

Go语言编程

威廉·肯尼迪

Publisher Resources

ISBN: 9787111562221