Partitioned Learned Bloom Filter

Partitioned Learned Bloom Filter
复制标题

DOI:
--
复制
发表时间:
2020-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Kapil Vaidya;Eric R. Knorr;Tim Kraska;M. Mitzenmacher
Kapil Vaidya;Eric R. Knorr;Tim Kraska;M. Mitzenmacher
中科院分区:
其他
文献类型:
--
作者:
Kapil Vaidya;Eric R. Knorr;Tim Kraska;M. Mitzenmacher

文献摘要

被引文献

相似文献

博学的Bloom过滤器通过为代表的数据集使用学习的模型来增强标准BLOOM过滤器。但是,学习的Bloom过滤器可能通过不充分利用输出来利用模型的限制。博学的布鲁姆过滤器通过简单地应用阈值来使用输出评分,其元素高于阈值被解释为阳性,而阈值以下的元素可独立于输出分数(使用较小的备用bloom滤波器来防止假阴性),以进一步分析(以进一步的分析) 。尽管最近的工作提出了其他启发式方法来更好地利用分数,但结果仅是启发式方法。在这里,我们相反,将最佳模型利用率作为优化问题构架。我们表明,可以有效地有效地解决了优化问题,从而产生了改进的{分区学习的Bloom Filter},该{分区的Bloom Filter}可以分配分数空间并利用每个区域的单独的备用Bloom过滤器。来自模拟数据集和现实世界数据集的实验结果表明,我们对原始学习的Bloom滤波器结构和先前提出的启发式改进的优化方法都有显着的性能提高。
Learned Bloom filters enhance standard Bloom filters by using a learned model for the represented data set. However, a learned Bloom filter may under-utilize the model by not taking full advantage of the output. The learned Bloom filter uses the output score by simply applying a threshold, with elements above the threshold being interpreted as positives, and elements below the threshold subject to further analysis independent of the output score (using a smaller backup Bloom filter to prevent false negatives). While recent work has suggested additional heuristic approaches to take better advantage of the score, the results are only heuristic. Here, we instead frame the problem of optimal model utilization as an optimization problem. We show that the optimization problem can be effectively solved efficiently, yielding an improved {partitioned learned Bloom filter}, which partitions the score space and utilizes separate backup Bloom filters for each region. Experimental results from both simulated and real-world datasets show significant performance improvements from our optimization approach over both the original learned Bloom filter constructions and previously proposed heuristic improvements.