An algorithm to identify medical practices common to both the General Practice Research Database and The Health Improvement Network database

An algorithm to identify medical practices common to both the General Practice Research Database and The Health Improvement Network database
复制标题

DOI:
10.1002/pds.3277
复制
发表时间:
2012-07-01
影响因子:
2.6
通讯作者:
Watson, Douglas J.
Watson, Douglas J.
中科院分区:
医学4区
文献类型:
--
作者:
Cai, Bing;Xu, Weifeng;Watson, Douglas J.

文献摘要

被引文献

相似文献

目的:识别全科医学研究数据库和健康改善网络数据库的共同实践,以便在没有重复记录的情况下合并数据库进行分析。方法:我们开发了两个独立的算法来识别两个数据库的共同做法。第一个使用治疗和临床数据集中的患者总数以及研究期间每年使用依托考昔和塞来考昔的患者总数。第二种方法使用按性别和四种不同出生年份分层的患者总数。通过按出生年份、临床访视日期和诊断代码比较患者级别的医疗记录,对两种算法识别的潜在匹配实践对进行进一步检查。结果两种算法共找到312个潜在匹配的实践对。另外15个潜在对仅通过一种算法匹配:13个仅通过算法1(A1)匹配,2个仅通过算法2(A2)匹配。检查患者级别的访视日期和诊断代码的匹配显示,所有327对潜在的重复实践实际上是两个数据库中的相同实践。结论这两种算法成功地找到了两个不同数据库的共同实践,而没有去识别实践。共同做法的识别允许结合两个数据库,没有重复的记录,以创建一个更大的数据集进行分析,与168个更多的做法时,单独使用的一般做法研究数据库,或与268个更多的做法时,单独使用健康改善网络。版权所有(c)2012约翰威利父子有限公司
Purpose To identify practices common to both the General Practice Research Database and The Health Improvement Network database for purposes of combining the databases for analysis without duplicate records. Methods We developed two independent algorithms to identify practices common to the two databases. The first used the total number of patients in the therapy and clinical data sets and the total number of etoricoxib and celecoxib users each year during the study period. The second used the total number of patients stratified by gender and four different categories of birth year. Further checking of potential matched practice pairs identified by the two algorithms was performed by comparing the patient-level medical records by birth year, dates of clinical visits, and diagnosis codes. Results Three hundred twelve potential matched pairs of practices were found by both algorithms. Fifteen additional potential pairs were matched by only one algorithm: 13 by algorithm 1 (A1) only and 2 by algorithm 2 (A2) only. The examination of the patient-level visit dates and diagnosis codes for the matches revealed that all of the 327 potential pairs of duplicate practices were in fact the same practice in the two databases. Conclusions The two algorithms successfully found the practices common to the two different databases without de-identifying the practices. The identification of the common practices allows for combining the two databases without duplicate records to create a larger data set for analysis, with 168 more practices than when using the General Practice Research Database alone, or with 268 more practices than when using The Health Improvement Network alone. Copyright (c) 2012 John Wiley & Sons, Ltd.