A Blocking Framework for Entity Resolution in Highly Heterogeneous Information Spaces

A Blocking Framework for Entity Resolution in Highly Heterogeneous Information Spaces
复制标题

DOI:
10.1109/tkde.2012.150
复制
发表时间:
2013-12
影响因子:
8.9
通讯作者:
G. Papadakis;Ekaterini Ioannou;Themis Palpanas;C. Niederée;W. Nejdl
G. Papadakis;Ekaterini Ioannou;Themis Palpanas;C. Niederée;W. Nejdl
中科院分区:
计算机科学2区
文献类型:
--
作者:
G. Papadakis;Ekaterini Ioannou;Themis Palpanas;C. Niederée;W. Nejdl

文献摘要

被引文献

相似文献

在高度异构的、嘈杂的、用户生成的实体集合中的实体解析(ER)的上下文中,几乎所有的块构建方法都采用冗余来实现高效率。然而,这种做法会导致大量的成对比较,对效率产生负面影响。现有的块处理策略旨在丢弃不必要的比较,而不会降低效率。在本文中,我们系统化的阻塞方法,清洁清洁ER(本质上是二次任务)通过由两个正交层组成的新型框架在高度异构的信息空间(HHIS)上进行:有效性层包含用于构建重叠块的方法,错过匹配的可能性很小;效率层包括显著限制成对比较的所需数量的多种技术,对检测到的重复的数量具有可控的影响。我们映射到我们的框架,所有相关的现有方法创建和处理块的上下文中HHIS,并提出了两种新的技术:属性聚类阻塞和比较调度。我们在两个大规模的真实数据集上评估了每个层和方法的性能,并验证了它们在效率和有效性之间的出色平衡。
In the context of entity resolution (ER) in highly heterogeneous, noisy, user-generated entity collections, practically all block building methods employ redundancy to achieve high effectiveness. This practice, however, results in a high number of pairwise comparisons, with a negative impact on efficiency. Existing block processing strategies aim at discarding unnecessary comparisons at no cost in effectiveness. In this paper, we systemize blocking methods for clean-clean ER (an inherently quadratic task) over highly heterogeneous information spaces (HHIS) through a novel framework that consists of two orthogonal layers: the effectiveness layer encompasses methods for building overlapping blocks with small likelihood of missed matches; the efficiency layer comprises a rich variety of techniques that significantly restrict the required number of pairwise comparisons, having a controllable impact on the number of detected duplicates. We map to our framework all relevant existing methods for creating and processing blocks in the context of HHIS, and additionally propose two novel techniques: attribute clustering blocking and comparison scheduling. We evaluate the performance of each layer and method on two large-scale, real-world data sets and validate the excellent balance between efficiency and effectiveness that they achieve.