Probabilistic record linkage and a method to calculate the positive predictive value

Probabilistic record linkage and a method to calculate the positive predictive value
复制标题

DOI:
10.1093/ije/31.6.1246
复制
发表时间:
2002-12-01
影响因子:
7.7
通讯作者:
Salmond, C
Salmond, C
中科院分区:
医学1区
文献类型:
--
作者:
Blakely, T;Salmond, C

文献摘要

被引文献

相似文献

背景计算机记录链接是常用的队列研究,以确定研究结果,因此其准确性分类的结果可以使用标准的流行病学术语的敏感性和阳性预测值(PPV)。方法我们描述了一个“重复的方法”来计算PPV的记录链接时,每个记录只能参与一个匹配(例如,人口文件连接到死亡文件)。该方法不需要验证两个文件中包含详细个人信息(例如姓名和地址)的记录子集,因此非常适合使用匿名数据的链接项目。重复方法假设一个文件中的记录数为0、1、2等,来自另一个文件的链接以组合概率预测的方式分布。做出这个假设后,假阳性链接的数量以及PPV是可以估计的。我们证明了这种重复的方法,使用输出匿名和概率记录联系的人口普查和死亡率records.Results的PPV估计符合预期的模式的基础上的概率记录联系的基本理论,敏感性分析和稳健。我们鼓励其他研究人员进一步评估这种方法的准确性。
Background Computerized record linkage is commonly used in cohort studies to ascertain the study outcome, and as such its accuracy classifying the outcome can be described using the standard epidemiological terms of sensitivity and positive predictive value (PPV).Method We describe a 'duplicate method' to calculate the PPV of record linkage when each record can only be involved in one match (e.g. linking population files to death files). The method does not require a validation subset of records from both files with detailed personal information (e.g. name and address), and is therefore ideal for linkage projects using anonymous data. The duplicate method assumes that the number of records from one file with zero, one, two, etc., links from the other file is distributed in a manner predicted by combinatorial probabilities. Having made this assumption, the number of false positive links, and hence the PPV, are estimable. We demonstrate this duplicate method using output from anonymous and probabilistic record linkage of census and mortality records in New Zealand.Results The PPV estimates conform to the pattern expected based on the underlying theory of probabilistic record linkage, and were robust to sensitivity analyses. We encourage other researchers to further assess the accuracy of this method.