Defining Disease Phenotypes Using National Linked Electronic Health Records: A Case Study of Atrial Fibrillation

Defining Disease Phenotypes Using National Linked Electronic Health Records: A Case Study of Atrial Fibrillation
复制标题

DOI:
10.1371/journal.pone.0110900
复制
发表时间:
2014-11-04
期刊:
影响因子:
3.7
通讯作者:
Hemingway, Harry
Hemingway, Harry
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Morley, Katherine I.;Wallace, Joshua;Hemingway, Harry

文献摘要

被引文献

相似文献

背景资料:国家电子健康记录(EHR)越来越多地用于研究,但由于来源之间捕获的信息存在差异,识别疾病病例具有挑战性。G.初级和二级保健)。我们的目标是提供一个透明的,可重复的模型,用于整合这些数据,使用心房颤动(AF),一种慢性疾病,在不同的医疗保健环境中以多种方式诊断和管理,作为一个案例study.Methods:在四个编码系统中确定了AF筛查,诊断和管理的潜在相关代码:阅读(初级保健诊断和程序),英国国家处方集(BNF;初级保健处方),ICD-10(二级保健诊断)和OPCS-4(二级保健程序)。从这些,我们开发了一个表型算法,通过专家审查和分析的链接EHR数据从1998年至2010年的队列214万英国患者年龄>= 30岁。队列也被用来评估表型通过检查事件AF和已知的危险factors.Results之间的关联:表型算法纳入286个代码:201读,63 BNF,18 ICD-10,和4 OPCS-4。记录了72,793例患者的AF诊断事件,但仅39.6%(N = 28,795)在初级保健和二级保健中记录。另外7 468例潜在病例是从治疗和先前存在的疾病数据中推断出来的。从每个来源确定的病例比例因诊断年龄而异;推断诊断在年轻病例中占更大比例(= 80岁)主要诊断为SC。(高血压、心肌梗死、心力衰竭)与使用不同EHR来源定义的AF事件的发生率在幅度上与传统的同意队列相当。单一的EHR来源不足以识别所有患者,也不能提供代表性样本。结合多个数据源,整合治疗和共病状况的信息,可以大大提高病例识别。
Background: National electronic health records (EHR) are increasingly used for research but identifying disease cases is challenging due to differences in information captured between sources (e. g. primary and secondary care). Our objective was to provide a transparent, reproducible model for integrating these data using atrial fibrillation (AF), a chronic condition diagnosed and managed in multiple ways in different healthcare settings, as a case study.Methods: Potentially relevant codes for AF screening, diagnosis, and management were identified in four coding systems: Read (primary care diagnoses and procedures), British National Formulary (BNF; primary care prescriptions), ICD-10 (secondary care diagnoses) and OPCS-4 (secondary care procedures). From these we developed a phenotype algorithm via expert review and analysis of linked EHR data from 1998 to 2010 for a cohort of 2.14 million UK patients aged >= 30 years. The cohort was also used to evaluate the phenotype by examining associations between incident AF and known risk factors.Results: The phenotype algorithm incorporated 286 codes: 201 Read, 63 BNF, 18 ICD-10, and four OPCS-4. Incident AF diagnoses were recorded for 72,793 patients, but only 39.6% (N = 28,795) were recorded in primary care and secondary care. An additional 7,468 potential cases were inferred from data on treatment and pre-existing conditions. The proportion of cases identified from each source differed by diagnosis age; inferred diagnoses contributed a greater proportion of younger cases (= 80 years) were mainly diagnosed in SC. Associations of risk factors (hypertension, myocardial infarction, heart failure) with incident AF defined using different EHR sources were comparable in magnitude to those from traditional consented cohorts.Conclusions: A single EHR source is not sufficient to identify all patients, nor will it provide a representative sample. Combining multiple data sources and integrating information on treatment and comorbid conditions can substantially improve case identification.