ProbClean: A probabilistic duplicate detection system

ProbClean: A probabilistic duplicate detection system
复制标题

ProbClean:概率重复检测系统

DOI:
--
复制
发表时间:
2010
期刊:
IEEE International Conference on Data Engineering
影响因子:
--
通讯作者:
Yubin Kim
Yubin Kim
中科院分区:
--
文献类型:
--
作者:
G. Beskales;Mohamed A. Soliman;Ihab F. Ilyas;Shai Ben;Yubin Kim

文献摘要

被引文献

相似文献

最突出的数据质量问题之一是存在重复记录。当前的数据清洗系统通常通过仔细选择重复检测算法的参数来产生输入数据的一个干净实例(修复)。找到正确的参数设置可能很难,在许多情况下,完美的设置并不存在。我们提出了ProbClean,一个系统,将重复检测程序作为具有不确定结果的数据处理任务。我们使用了一种新的不确定性模型,该模型对对应于不同参数设置的可能修复空间进行了压缩编码。ProbClean有效地支持关系查询,并允许针对一组可能的修复进行新类型的查询。
One of the most prominent data quality problems is the existence of duplicate records. Current data cleaning systems usually produce one clean instance (repair) of the input data, by carefully choosing the parameters of the duplicate detection algorithms. Finding the right parameter settings can be hard, and in many cases, perfect settings do not exist. We propose ProbClean, a system that treats duplicate detection procedures as data processing tasks with uncertain outcomes. We use a novel uncertainty model that compactly encodes the space of possible repairs corresponding to different parameter settings. ProbClean efficiently supports relational queries and allows new types of queries against a set of possible repairs.