Efficient similarity joins for near duplicate detection

Efficient similarity joins for near duplicate detection
复制标题

DOI:
10.1145/1367497.1367516
复制
发表时间:
2008-04
期刊:
--
影响因子:
--
通讯作者:
Chuan Xiao;Wei Wang;Xuemin Lin;J. Yu;Guoren Wang
Chuan Xiao;Wei Wang;Xuemin Lin;J. Yu;Guoren Wang
中科院分区:
其他
文献类型:
--
作者:
Chuan Xiao;Wei Wang;Xuemin Lin;J. Yu;Guoren Wang

文献摘要

被引文献

相似文献

随着数据量的不断增加以及需要集成来自多个数据源的数据,一个具有挑战性的问题是有效地找到近似重复的记录。在本文中,我们专注于高效的算法来找到对记录,使他们的相似性是在给定的阈值以上。几个现有的算法依赖于前缀过滤原则,以避免计算所有可能的记录对的相似性值。我们提出了新的过滤技术,利用排序信息,他们被集成到现有的方法,大大减少了候选人的大小,从而提高效率。实验结果表明,在多个真实的数据集上,我们提出的算法可以实现2.6 - 5倍的加速,并为近似重复网页检测问题提供了替代解决方案。
With the increasing amount of data and the need to integrate data from multiple data sources, a challenging issue is to find near duplicate records efficiently. In this paper, we focus on efficient algorithms to find pairs of records such that their similarities are above a given threshold. Several existing algorithms rely on the prefix filtering principle to avoid computing similarity values for all possible pairs of records. We propose new filtering techniques by exploiting the ordering information; they are integrated into the existing methods and drastically reduce the candidate sizes and hence improve the efficiency. Experimental results show that our proposed algorithms can achieve up to 2.6x - 5x speed-up over previous algorithms on several real datasets and provide alternative solutions to the near duplicate Web page detection problem.