Comparative Analysis of Approximate Blocking Techniques for Entity Resolution

Comparative Analysis of Approximate Blocking Techniques for Entity Resolution
复制标题

DOI:
10.14778/2947618.2947624
复制
发表时间:
2016-05
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
G. Papadakis;Jonathan Svirsky;A. Gal;Themis Palpanas
G. Papadakis;Jonathan Svirsky;A. Gal;Themis Palpanas
中科院分区:
其他
文献类型:
--
作者:
G. Papadakis;Jonathan Svirsky;A. Gal;Themis Palpanas

文献摘要

被引文献

相似文献

实体解析是数据集合合并的核心任务。由于其二次复杂度,它通常通过阻塞扩展到大量数据:相似的实体聚集到块中,并且仅在共同发生的实体之间执行成对比较,代价是丢失一些匹配。有许多阻断方法,这项工作的目的是提供一个全面的实证调查,扩展比较的维度超出了什么是在文献中普遍可用。我们考虑了17种最先进的阻塞方法,并使用6个流行的真实数据集来检查其内部配置的鲁棒性以及它们在有效性和时间效率之间的相对平衡。我们还研究了它们在7个已建立的合成数据集的语料库上的可扩展性,这些数据集的范围从10,000到200万个实体。
Entity Resolution is a core task for merging data collections. Due to its quadratic complexity, it typically scales to large volumes of data through blocking: similar entities are clustered into blocks and pair-wise comparisons are executed only between co-occurring entities, at the cost of some missed matches. There are numerous blocking methods, and the aim of this work is to offer a comprehensive empirical survey, extending the dimensions of comparison beyond what is commonly available in the literature. We consider 17 state-of-the-art blocking methods and use 6 popular real datasets to examine the robustness of their internal configurations and their relative balance between effectiveness and time efficiency. We also investigate their scalability over a corpus of 7 established synthetic datasets that range from 10,000 to 2 million entities.