An efficient record linkage scheme using graphical analysis for identifier error detection

An efficient record linkage scheme using graphical analysis for identifier error detection
复制标题

DOI:
10.1186/1472-6947-11-7
复制
发表时间:
2011-02-01
影响因子:
3.5
通讯作者:
Wyllie, David H.
Wyllie, David H.
中科院分区:
医学3区
文献类型:
--
作者:
Finney, John M.;Walker, A. Sarah;Wyllie, David H.

文献摘要

被引文献

相似文献

背景资料:个人信息的整合(记录链接)是医疗服务、流行病学和“商业智能”应用中的一个关键问题。它现在是常见的是需要链接非常大量的记录,往往包含各种组合的理论上唯一的标识符,如NHS号码,这是不完整的和容易出错的。方法:我们描述了一个两步的记录链接算法,其中具有高基数的标识符被识别或生成,并用于执行一个初始的基于精确匹配的链接。随后,所产生的集群进行了研究,如果合适的话,分区使用基于图的算法检测错误identifiers.Results:该系统被用来聚类超过2.5亿健康记录从5个数据源在一个大的英国医院集团。在大约30分钟内完成的链接产生了360万个集群,其中约99.8%包含来自一个患者的记录,并且可能性很高。虽然计算效率高,算法的要求,每个记录的至少一个标识符的精确匹配到另一个集群formation.Conclusions的可能是一个限制,在一些数据库中包含低标识符的记录quality.Conclusions:所描述的技术提供了一个简单,快速和高效的两步方法,大规模的初始链接记录通常发现在英国的国家卫生服务。
Background: Integration of information on individuals (record linkage) is a key problem in healthcare delivery, epidemiology, and "business intelligence" applications. It is now common to be required to link very large numbers of records, often containing various combinations of theoretically unique identifiers, such as NHS numbers, which are both incomplete and error-prone.Methods: We describe a two-step record linkage algorithm in which identifiers with high cardinality are identified or generated, and used to perform an initial exact match based linkage. Subsequently, the resulting clusters are studied and, if appropriate, partitioned using a graph based algorithm detecting erroneous identifiers.Results: The system was used to cluster over 250 million health records from five data sources within a large UK hospital group. Linkage, which was completed in about 30 minutes, yielded 3.6 million clusters of which about 99.8% contain, with high likelihood, records from one patient. Although computationally efficient, the algorithm's requirement for exact matching of at least one identifier of each record to another for cluster formation may be a limitation in some databases containing records of low identifier quality.Conclusions: The technique described offers a simple, fast and highly efficient two-step method for large scale initial linkage for records commonly found in the UK's National Health Service.