A Domain-Independent Data Cleaning Algorithm for Detecting Similar-Duplicates

A Domain-Independent Data Cleaning Algorithm for Detecting Similar-Duplicates
复制标题

用于检测相似重复项的与域无关的数据清理算法

DOI:
--
复制
发表时间:
2010
期刊:
Journal of Computers
影响因子:
--
通讯作者:
G. Rahaman
G. Rahaman
中科院分区:
--
文献类型:
--
作者:
Kazi Shah Nawaz Ripon;A. F. M. Moshiur Rahman;G. Rahaman

文献摘要

被引文献

相似文献

数据挖掘算法通常假设数据是干净一致的。然而,在实践中,情况并非总是如此,因此,检测和消除重复记录是数据清理的重要组成部分。类似重复记录的存在会导致数据的过度表示。如果数据库包含相同数据的不同表示,则从数据挖掘算法获得的结果将是错误的。检测相似重复的记录是一项困难的任务,特别是当记录是域独立的时候。在本文中,我们提出了一种新的域无关技术来更好地协调相似重复记录。我们还介绍了使相似重复检测算法更快、更有效的新思路。此外,本文还对及物性规则进行了重大修改。最后,我们提出了一种算法,该算法将所有这些相似重复检测技术整合到一个领域独立的环境中。将所提方法的性能与其他方法进行了比较,实验结果证实了所提方法的优越性。
Normal 0 MicrosoftInternetExplorer4 Data mining algorithms generally assume that data will be clean and consistent. However, in practice, this is not always the case, and for this reason the detection and elimination of duplicate records is an important part of data cleaning. The presence of similar-duplicate records causes over-representation of data. If the database contains different representations of the same data, the results obtained from the data mining algorithm will be erroneous. The detection of similar-duplicate records is a difficult task, especially when the records are domain-independent. In this paper, we propose a novel domain-independent technique for better reconciling the similar-duplicate records. We also introduce new ideas for making similar-duplicate detection algorithms faster and more efficient. In addition, a significant modification of the transitivity rule is also proposed. Finally, we propose an algorithm that incorporates all these techniques for similar-duplicate detection into a domain-independent environment. The performance of the proposed method has been compared to other methods and the superiority of the proposed method has been confirmed by the experimental results.