The merge/purge problem for large databases

The merge/purge problem for large databases
复制标题

DOI:
10.1145/223784.223807
复制
发表时间:
1995-05
期刊:
--
影响因子:
--
通讯作者:
Mauricio A. Hernández;S. Stolfo
Mauricio A. Hernández;S. Stolfo
中科院分区:
其他
文献类型:
--
作者:
Mauricio A. Hernández;S. Stolfo

文献摘要

被引文献

相似文献

许多商业组织经常收集大量的数据库,用于各种营销和业务分析功能。任务是通过识别出现在许多不同数据库中的不同个体来关联来自不同数据库的信息,这些个体通常以不一致且通常不正确的方式出现。我们在这里研究的问题是以尽可能有效的方式合并来自多个来源的数据,同时最大限度地提高结果的准确性。我们称之为合并/清除问题。在本文中,我们详细介绍了排序的邻域方法,一些用来解决合并/清除,目前的实验结果表明,这种方法可能在实践中工作得很好,但在很大的代价。基于聚类的另一种方法也提出了一个比较评价排序邻域方法。我们展示了一种基于多遍方法的提高结果准确性的方法,该方法通过在每次传递中考虑替代主键属性的独立运行结果上计算传递闭包来成功。
Many commercial organizations routinely gather large numbers of databases for various marketing and business analysis functions. The task is to correlate information from different databases by identifying distinct individuals that appear in a number of different databases typically in an inconsistent and often incorrect fashion. The problem we study here is the task of merging data from multiple sources in as efficient manner as possible, while maximizing the accuracy of the result. We call this the merge/purge problem. In this paper we detail the sorted neighborhood method that is used by some to solve merge/purge and present experimental results that demonstrates this approach may work well in practice but at great expense. An alternative method based upon clustering is also presented with a comparative evaluation to the sorted neighborhood method. We show a means of improving the accuracy of the results based upon a multi-pass approach that succeeds by computing the Transitive Closure over the results of independent runs considering alternative primary key attributes in each pass.