Accuracy of Probabilistic Linkage Using the Enhanced Matching System for Public Health and Epidemiological Studies.

Accuracy of Probabilistic Linkage Using the Enhanced Matching System for Public Health and Epidemiological Studies.
复制标题

DOI:
10.1371/journal.pone.0136179
复制
发表时间:
2015
期刊:
影响因子:
3.7
通讯作者:
Abubakar I
Abubakar I
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Aldridge RW;Shaji K;Hayward AC;Abubakar I

文献摘要

被引文献

相似文献

增强匹配系统(EMS)是英国公共卫生结核病部门开发的一种概率记录链接程序,用于匹配两个数据集上的个人数据。本文概述了EMS是如何工作的,并调查了它在公共卫生数据集之间联系的准确性。EMS是一个可配置的Microsoft SQL Server数据库程序。为了检查EMS的准确性,使用国民健康服务(NHS)号码作为黄金标准唯一标识符对两个公共卫生数据库进行了匹配。然后,在不包括NHS编号的情况下,对相同的两个数据集执行概率链接。进行了灵敏度分析,考察了不同匹配工艺参数的影响。在两个数据集(包含5931条和1759条记录)之间使用NHS编号进行精确匹配,确定了1071对匹配对。EMS概率连锁鉴定出1068个记录对。概率连锁的敏感性为99.5%(95%CI:98.9,99.8),特异性为100.0%(95%CI:99.9,100.0),阳性预测值为99.8%(95%CI:99.3,100.0),阴性预测值为99.9%(95%CI:99.8,100.0)。当包含地址变量并使用自动生成的阈值通过手动审查确定链接时,概率匹配最准确。随着卫生和社会保健领域的国家电子数据集的建立,EMS使以前无法回答的研究问题能够自信地得到解决,并对联系过程的准确性充满信心。在将小样本匹配到非常大的数据库(如全国医院就诊记录)的情况下,与此分析中提供的结果相比,阳性预测值或灵敏度可能会根据数据库之间匹配的流行率而下降。尽管有这一可能的限制,但在无法使用共同识别符进行精确匹配的情况下,包括在低收入环境中,以及对于无家可归人口等弱势群体而言,概率联系具有很大的潜力,在这些群体中,缺乏唯一识别符和较低的数据质量历来阻碍了在数据集中识别个人的能力。
The Enhanced Matching System (EMS) is a probabilistic record linkage program developed by the tuberculosis section at Public Health England to match data for individuals across two datasets. This paper outlines how EMS works and investigates its accuracy for linkage across public health datasets. EMS is a configurable Microsoft SQL Server database program. To examine the accuracy of EMS, two public health databases were matched using National Health Service (NHS) numbers as a gold standard unique identifier. Probabilistic linkage was then performed on the same two datasets without inclusion of NHS number. Sensitivity analyses were carried out to examine the effect of varying matching process parameters. Exact matching using NHS number between two datasets (containing 5931 and 1759 records) identified 1071 matched pairs. EMS probabilistic linkage identified 1068 record pairs. The sensitivity of probabilistic linkage was calculated as 99.5% (95%CI: 98.9, 99.8), specificity 100.0% (95%CI: 99.9, 100.0), positive predictive value 99.8% (95%CI: 99.3, 100.0), and negative predictive value 99.9% (95%CI: 99.8, 100.0). Probabilistic matching was most accurate when including address variables and using the automatically generated threshold for determining links with manual review. With the establishment of national electronic datasets across health and social care, EMS enables previously unanswerable research questions to be tackled with confidence in the accuracy of the linkage process. In scenarios where a small sample is being matched into a very large database (such as national records of hospital attendance) then, compared to results presented in this analysis, the positive predictive value or sensitivity may drop according to the prevalence of matches between databases. Despite this possible limitation, probabilistic linkage has great potential to be used where exact matching using a common identifier is not possible, including in low-income settings, and for vulnerable groups such as homeless populations, where the absence of unique identifiers and lower data quality has historically hindered the ability to identify individuals across datasets.