A Graph-Based Blocking Approach for Entity Matching Using Contrastively Learned Embeddings

A Graph-Based Blocking Approach for Entity Matching Using Contrastively Learned Embeddings
复制标题

DOI:
10.1145/3584014.3584017
复制
发表时间:
2022-12
期刊:
ACM SIGAPP Applied Computing Review
影响因子:
--
通讯作者:
John Bosco Mugeni;Toshiyuki Amagasa
John Bosco Mugeni;Toshiyuki Amagasa
中科院分区:
其他
文献类型:
--
作者:
John Bosco Mugeni;Toshiyuki Amagasa

文献摘要

相似文献

数据集成被认为是实体匹配过程中的一项重要任务。在这个过程中,必须识别和消除冗余和狡猾的条目,以提高数据质量。为了将其存档,执行所有实体之间的比较。然而,这具有二次计算复杂度。为了避免这种情况,“阻塞”将比较限制在可能的匹配上。本文提出了一种基于k近邻图的分块方法,该方法利用了来自预训练transformers的最先进的上下文感知句子嵌入。我们的方法将每个数据库元组映射到一个节点,并生成一个图,其中节点通过边连接,如果它们相关。然后,我们调用无监督的社区检测技术在这个图和治疗阻塞作为一个图聚类问题。我们的工作的动机是在现实世界中的场景中的实体匹配的训练数据的稀缺性和有限的可扩展性的阻塞方案中存在的激增的数据。此外,我们研究了对比训练的嵌入对上述系统的影响,并在四个数据集上测试了其能力,这些数据集显示了超过600万次比较。我们表明,由于k-最近邻图的有效数据结构,我们在目标基准上的块处理时间各不相同。我们的研究结果还表明,与当前基于深度学习的阻塞解决方案相比,我们的方法在F1得分方面具有更好的性能。
Data integration is considered a crucial task in the entity matching process. In this process, redundant and cunning entries must be identified and eliminated to improve the data quality. To archive this, a comparison between all entities is performed. However, this has quadratic computational complexity. To avoid this, `blocking' limits comparisons to probable matches. This paper presents a k-nearest neighbor graph-based blocking approach utilizing state-of-the-art context-aware sentence embeddings from pre-trained transformers. Our approach maps each database tuple to a node and generates a graph where nodes are connected by edges if they are related. We then invoke unsupervised community detection techniques over this graph and treat blocking as a graph clustering problem. Our work is motivated by the scarcity of training data for entity matching in real-world scenarios and the limited scalability of blocking schemes in the presence of proliferating data. Additionally, we investigate the impact of contrastively trained embeddings on the above system and test its capabilities on four data sets exhibiting more than 6 million comparisons. We show that our block processing times on the target benchmarks vary owing to the efficient data structure of the k-nearest neighbor graph. Our results also show that our method achieves better performance in terms of F1 score when compared to current deep learning-based blocking solutions.