Duplicate Record Detection: A Survey

Duplicate Record Detection: A Survey
复制标题

DOI:
10.1109/tkde.2007.9
复制
发表时间:
2007
影响因子:
8.9
通讯作者:
A. Elmagarmid;Panagiotis G. Ipeirotis;V. Verykios
A. Elmagarmid;Panagiotis G. Ipeirotis;V. Verykios
中科院分区:
计算机科学2区
文献类型:
--
作者:
A. Elmagarmid;Panagiotis G. Ipeirotis;V. Verykios

文献摘要

被引文献

相似文献

在现实世界中,实体通常在数据库中有两个或多个表示形式。重复的记录不共享一个公共键和/或它们包含错误,使重复匹配成为一项困难的任务。错误是由于转录错误、信息不完整、缺乏标准格式或这些因素的任何组合而引入的。在本文中,我们对重复记录检测的文献进行了全面的分析。我们讨论了通常用于检测相似字段条目的相似性度量,并提出了一套广泛的重复检测算法,可以检测数据库中近似重复的记录。我们还介绍了用于提高近似重复检测算法的效率和可扩展性的多种技术。最后,我们对现有工具进行了介绍,并简要讨论了该领域的重大开放问题
Often, in the real world, entities have two or more representations in databases. Duplicate records do not share a common key and/or they contain errors that make duplicate matching a difficult task. Errors are introduced as the result of transcription errors, incomplete information, lack of standard formats, or any combination of these factors. In this paper, we present a thorough analysis of the literature on duplicate record detection. We cover similarity metrics that are commonly used to detect similar field entries, and we present an extensive set of duplicate detection algorithms that can detect approximately duplicate records in a database. We also cover multiple techniques for improving the efficiency and scalability of approximate duplicate detection algorithms. We conclude with coverage of existing tools and with a brief discussion of the big open problems in the area