A Generalized Fellegi-Sunter Framework for Multiple Record Linkage With Application to Homicide Record Systems

A Generalized Fellegi-Sunter Framework for Multiple Record Linkage With Application to Homicide Record Systems
复制标题

DOI:
10.1080/01621459.2012.757231
复制
发表时间:
2013-06-01
影响因子:
3.7
通讯作者:
Fienberg, Stephen E.
Fienberg, Stephen E.
中科院分区:
数学1区
文献类型:
--
作者:
Sadinle, Mauricio;Fienberg, Stephen E.

文献摘要

被引文献

相似文献

我们提出了一个概率方法连接多个文件。在没有所记录的个人的独特识别资料的情况下,这项任务并非微不足道。这是将普查数据与覆盖面测量调查联系起来进行普查覆盖面评价时的常见情况,一般来说,这是需要整合多个记录系统进行事后分析时的常见情况。我们的方法概括了Fellegi-Sunter理论,用于连接来自两个文件夹的记录及其调制解调器实现。多记录连接的目标是根据不同的匹配模式对来自K个记录的记录K元组进行分类。我们的方法采用了传递性的协议在用于模型匹配概率的数据的计算。我们使用混合模型来拟合匹配概率通过最大似然使用期望最大化算法。我们提出了一种方法来决定记录的K-元组的匹配模式的子集的成员资格,我们证明了它的最优性。我们将我们的方法应用到三个哥伦比亚凶杀案记录系统的集成中,并进行了模拟研究,以探索测量误差和不同场景下该方法的性能。所提出的方法效果良好,并为未来的研究开辟了新的方向。
We present a probabilistic method for linking multiple datafiles. This task is not trivial in the absence of unique identifiers for the individuals recorded. This is a common scenario when linking census data to coverage measurement surveys for census coverage evaluation, and in general when multiple record systems need to be integrated for posterior analysis. Our method generalizes the Fellegi-Sunter theory for linking records from two datafiles and its modem implementations. The goal of multiple record linkage is to classify the record K-tuples coming from K datafiles according to the different matching patterns. Our method incorporates the transitivity of agreement in the computation of the data used to model matching probabilities. We use a mixture model to fit matching probabilities via maximum likelihood using the Expectation-Maximization algorithm. We present a method to decide the record K-tuples membership to the subsets of matching patterns and we prove its optimality. We apply our method to the integration of the three Colombian homicide record systems and perform a simulation study to explore the performance of the method under measurement error and different scenarios. The proposed method works well and opens new directions for future research.