Multifile Partitioning for Record Linkage and Duplicate Detection

Multifile Partitioning for Record Linkage and Duplicate Detection
复制标题

DOI:
10.1080/01621459.2021.2013242
复制
发表时间:
2021-10
影响因子:
3.7
通讯作者:
Serge Aleshin-Guendel;Mauricio Sadinle
Serge Aleshin-Guendel;Mauricio Sadinle
中科院分区:
数学1区
文献类型:
--
作者:
Serge Aleshin-Guendel;Mauricio Sadinle

文献摘要

相似文献

摘要在没有唯一标识符的情况下,合并包含关于重叠实体集的信息的文件是一项具有挑战性的任务,并且当一些实体在文件中重复时,合并文件变得更加复杂。解决这个问题的大多数方法都集中在链接两个假定没有重复的文件,或者检测单个文件中的哪些记录是重复的。然而,在实践中经常会遇到介于这两种设置之间或之外的场景。我们提出了一个贝叶斯方法的一般设置的多文件记录链接和重复检测。我们使用一种新的分区表示提出了一个结构化的分区,可以将先验信息的数据收集过程中的fixiles以灵活的方式,并扩展以前的模型比较数据,以适应多文件设置。我们还引入了一个家庭的损失函数,以获得贝叶斯估计的分区,允许不确定的部分的分区被遗留未解决的。我们提出的方法的性能进行了探讨,通过广泛的模拟。本文的补充材料可在网上查阅。
Abstract Merging datafiles containing information on overlapping sets of entities is a challenging task in the absence of unique identifiers, and is further complicated when some entities are duplicated in the datafiles. Most approaches to this problem have focused on linking two files assumed to be free of duplicates, or on detecting which records in a single file are duplicates. However, it is common in practice to encounter scenarios that fit somewhere in between or beyond these two settings. We propose a Bayesian approach for the general setting of multifile record linkage and duplicate detection. We use a novel partition representation to propose a structured prior for partitions that can incorporate prior information about the data collection processes of the datafiles in a flexible manner, and extend previous models for comparison data to accommodate the multifile setting. We also introduce a family of loss functions to derive Bayes estimates of partitions that allow uncertain portions of the partitions to be left unresolved. The performance of our proposed methodology is explored through extensive simulations. Supplementary materials for this article are available online.