BUBBLE : A Quality-Aware Human-in-the-loop Entity Matching Framework

BUBBLE : A Quality-Aware Human-in-the-loop Entity Matching Framework
复制标题

DOI:
10.1109/bigdata52589.2021.9672002
复制
发表时间:
2021-12
期刊:
2021 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Naofumi Osawa;Hiroyoshi Ito;Yukihiro Fukushima;Takashi Harada;Atsuyuki Morishima
Naofumi Osawa;Hiroyoshi Ito;Yukihiro Fukushima;Takashi Harada;Atsuyuki Morishima
中科院分区:
其他
文献类型:
--
作者:
Naofumi Osawa;Hiroyoshi Ito;Yukihiro Fukushima;Takashi Harada;Atsuyuki Morishima

文献摘要

相似文献

实体匹配是信息集成和数据清理中的一个重要问题。由于同一实体的表示形式各不相同,因此通常不可能完全自动化实体匹配并需要人工输入。然而,为了保证高质量的实体匹配,如何将人力资源整合到实体匹配中,同时使人力资源成本最小化?本文提出了一种融合贝叶斯推理和众包的新型人在环实体匹配框架BUBBLE。为了保证实体匹配的质量,通过贝叶斯推理来判断匹配是否需要众包。我们证明了我们可以为这个问题定义贝叶斯错误率。在优化方面,我们利用度量学习在学习的嵌入空间中通过最近邻搜索来选择候选匹配对,并构造一个k近邻图来避免冗余匹配。我们将BUBBLE应用于国会图书馆的书目数据匹配问题。实验结果表明,与相同数量的任务分配相比,BUBBLE可以为人类分配任务并获得更高质量的结果。结果还表明,该优化方案在不牺牲质量的前提下是有效的。
Entity matching is an issue of interest in information integration and data cleaning. Since the representations of the same entity vary, it is often impossible to fully automate the entity matching and require human inputs. However, to guarantee high-quality entity matching, how to integrate human resources into the entity matching while minimizing the cost of human resources? In this paper, we propose BUBBLE, a novel human-in-the-loop entity matching framework hybridizing Bayesian inference and crowdsourcing. To guarantee entity matching quality, Bayesian inference is conducted to determine whether the matching requires crowdsourcing. We show that we can define Bayesian error rate for this problem. For optimization, we use metric learning to select the candidate matching pairs by nearest-neighbor search in the learned embedding space, and we construct a k-nearest neighbor graph to avoid the redundant matching. We applied BUBBLE to a bibliographic data matching problem on the National Diet Library. The experimental results show that BUBBLE can assign tasks to humans with higher quality results compared to those of the same number of task assignments to humans. The result also shows that our optimization scheme is effective without sacrificing the quality.