CrowdER: Crowdsourcing Entity Resolution

CrowdER: Crowdsourcing Entity Resolution
复制标题

DOI:
10.14778/2350229.2350263
复制
发表时间:
2012-07
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Jiannan Wang;Tim Kraska;M. Franklin;Jianhua Feng
Jiannan Wang;Tim Kraska;M. Franklin;Jianhua Feng
中科院分区:
其他
文献类型:
--
作者:
Jiannan Wang;Tim Kraska;M. Franklin;Jianhua Feng

文献摘要

被引文献

相似文献

实体解析是数据集成和数据清理的核心。算法方法的质量一直在提高,但还远远不够完美。众包平台提供了一种更准确、但昂贵(且缓慢)的方式,让人类洞察这一过程。以前的工作已经提出了将批处理验证任务呈现给人类工作人员,但即使是批处理,由于需要测试大量匹配,对于中等规模的数据集,仅使用人工方法也是不可行的。相反,我们提出了一种混合的人机方法,在这种方法中,机器用于对所有数据进行初始的、粗略的传递,而人则用于仅验证最可能匹配的对。我们表明,对于这样的混合系统,生成给定大小的最小数量的验证任务是NP-Hard的,但我们开发了一种新的两层启发式方法来创建批处理任务。我们描述了这种方法,并展示了使用流行的众包平台在真实数据集上进行的大量实验结果。实验表明,与机器或人工混合方法相比,我们的混合方法具有较高的效率和精度。
Entity resolution is central to data integration and data cleaning. Algorithmic approaches have been improving in quality, but remain far from perfect. Crowdsourcing platforms offer a more accurate but expensive (and slow) way to bring human insight into the process. Previous work has proposed batching verification tasks for presentation to human workers but even with batching, a human-only approach is infeasible for data sets of even moderate size, due to the large numbers of matches to be tested. Instead, we propose a hybrid human-machine approach in which machines are used to do an initial, coarse pass over all the data, and people are used to verify only the most likely matching pairs. We show that for such a hybrid system, generating the minimum number of verification tasks of a given size is NP-Hard, but we develop a novel two-tiered heuristic approach for creating batched tasks. We describe this method, and present the results of extensive experiments on real data sets using a popular crowdsourcing platform. The experiments show that our hybrid approach achieves both good efficiency and high accuracy compared to machine-only or human-only alternatives.