Methods to Develop an Electronic Medical Record Phenotype Algorithm to Compare the Risk of Coronary Artery Disease across 3 Chronic Disease Cohorts.

Methods to Develop an Electronic Medical Record Phenotype Algorithm to Compare the Risk of Coronary Artery Disease across 3 Chronic Disease Cohorts.
复制标题

DOI:
10.1371/journal.pone.0136651
复制
发表时间:
2015
期刊:
影响因子:
3.7
通讯作者:
Cai T
Cai T
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Liao KP;Ananthakrishnan AN;Kumar V;Xia Z;Cagan A;Gainer VS;Goryachev S;Chen P;Savova GK;Agniel D;Churchill S;Lee J;Murphy SN;Plenge RM;Szolovits P;Kohane I;Shaw SY;Karlson EW;Cai T

文献摘要

被引文献

相似文献

通常,开发了使用电子病历(EMR)数据对表型进行分类的算法,以在特定患者群体中表现良好。人们对分析越来越感兴趣,这些分析可以允许研究不同疾病的特定结果。在EMR中进行这样的研究需要一种可以应用于不同患者人群的算法。我们的目标是:(1)开发一种算法,可以在不同的患者人群中研究冠状动脉疾病(CAD);(2)研究在算法中添加使用自然语言处理(NLP)提取的叙述数据的影响。此外,我们还演示了如何在初步研究中实现CAD算法来比较3种慢性疾病的风险。我们研究了来自两个大型学术中心的3个已建立的基于EMR的患者队列:糖尿病(DM,n = 65,099)、炎症性肠病(IBD,n = 10,974)和类风湿性关节炎(RA,n = 4,453)。我们在RA队列中使用NLP和结构化数据(例如ICD 9代码)开发了一种CAD算法,并在DM和IBD队列中进行了验证。除了结构化数据之外,使用NLP的CAD算法在训练集(RA)和验证集(IBD和DM)中实现了特异性>95%,阳性预测值(PPV)90%。增加NLP数据提高了所有队列的敏感性,将另外17%的CAD受试者分类为IBD和10%的DM,同时保持PPV为90%。该算法将16,488例DM(26.1%)、457例IBD(4.2%)和245例RA(5.0%)分类为CAD。在一项横断面分析中,在调整传统的心血管危险因素后,与DM相比,RA和IBD的CAD风险分别低63%和68%(p<0.0001)。我们开发并验证了一种CAD算法,该算法在不同的患者人群中表现良好。将NLP添加到CAD算法中提高了算法的灵敏度,特别是在CAD患病率较低的队列中。初步数据表明,与DM相比,RA和IBD的CAD风险显著降低。
Typically, algorithms to classify phenotypes using electronic medical record (EMR) data were developed to perform well in a specific patient population. There is increasing interest in analyses which can allow study of a specific outcome across different diseases. Such a study in the EMR would require an algorithm that can be applied across different patient populations. Our objectives were: (1) to develop an algorithm that would enable the study of coronary artery disease (CAD) across diverse patient populations; (2) to study the impact of adding narrative data extracted using natural language processing (NLP) in the algorithm. Additionally, we demonstrate how to implement CAD algorithm to compare risk across 3 chronic diseases in a preliminary study. We studied 3 established EMR based patient cohorts: diabetes mellitus (DM, n = 65,099), inflammatory bowel disease (IBD, n = 10,974), and rheumatoid arthritis (RA, n = 4,453) from two large academic centers. We developed a CAD algorithm using NLP in addition to structured data (e.g. ICD9 codes) in the RA cohort and validated it in the DM and IBD cohorts. The CAD algorithm using NLP in addition to structured data achieved specificity >95% with a positive predictive value (PPV) 90% in the training (RA) and validation sets (IBD and DM). The addition of NLP data improved the sensitivity for all cohorts, classifying an additional 17% of CAD subjects in IBD and 10% in DM while maintaining PPV of 90%. The algorithm classified 16,488 DM (26.1%), 457 IBD (4.2%), and 245 RA (5.0%) with CAD. In a cross-sectional analysis, CAD risk was 63% lower in RA and 68% lower in IBD compared to DM (p<0.0001) after adjusting for traditional cardiovascular risk factors. We developed and validated a CAD algorithm that performed well across diverse patient populations. The addition of NLP into the CAD algorithm improved the sensitivity of the algorithm, particularly in cohorts where the prevalence of CAD was low. Preliminary data suggest that CAD risk was significantly lower in RA and IBD compared to DM.