Development and Validation of an Algorithm to Identify Nonalcoholic Fatty Liver Disease in the Electronic Medical Record

Development and Validation of an Algorithm to Identify Nonalcoholic Fatty Liver Disease in the Electronic Medical Record
复制标题

DOI:
10.1007/s10620-015-3952-x
复制
发表时间:
2016-03-01
影响因子:
3.1
通讯作者:
Shaw, Stanley Y.
Shaw, Stanley Y.
中科院分区:
医学3区
文献类型:
--
作者:
Corey, Kathleen E.;Kartoun, Uri;Shaw, Stanley Y.

文献摘要

被引文献

相似文献

非酒精性脂肪肝(NAFLD)是全球慢性肝病最常见的原因。由于缺乏计算识别方法,NAFLD疾病进展和肝脏相关结局的风险因素仍不完全清楚。本研究旨在设计一种电子病历(EMR)中NAFLD的分类算法,用于大规模纵向队列的发展。从Partners Healthcare的研究患者数据登记处随机选择了620名患者的训练集。为了评估NAFLD的真正诊断,我们进行了病历审查,并考虑了活检或临床诊断的NAFLD文件。我们在模型变量中包括实验室测量值、诊断代码和从医疗记录中提取的概念。多变量分析中包括P < 0.05的变量,NAFLD分类算法包括EMR中自然语言提及NAFLD的次数、NAFLD的ICD-9编码的终生次数和甘油三酯水平。该分类算法上级单独使用ICD-9数据的算法,AUC为0.85对0.75(P < 0.0001),并导致创建了具有高NAFLD概率的8458个个体的新的独立队列。这种方法易于开发和部署,可以应用于不同的机构,以创建基于EMR的NAFLD患者队列。
Nonalcoholic fatty liver disease (NAFLD) is the most common cause of chronic liver disease worldwide. Risk factors for NAFLD disease progression and liver-related outcomes remain incompletely understood due to the lack of computational identification methods. The present study sought to design a classification algorithm for NAFLD within the electronic medical record (EMR) for the development of large-scale longitudinal cohorts.We implemented feature selection using logistic regression with adaptive LASSO. A training set of 620 patients was randomly selected from the Research Patient Data Registry at Partners Healthcare. To assess a true diagnosis for NAFLD we performed chart reviews and considered either a documentation of a biopsy or a clinical diagnosis of NAFLD. We included in our model variables laboratory measurements, diagnosis codes, and concepts extracted from medical notes. Variables with P < 0.05 were included in the multivariable analysis.The NAFLD classification algorithm included number of natural language mentions of NAFLD in the EMR, lifetime number of ICD-9 codes for NAFLD, and triglyceride level. This classification algorithm was superior to an algorithm using ICD-9 data alone with AUC of 0.85 versus 0.75 (P < 0.0001) and leads to the creation of a new independent cohort of 8458 individuals with a high probability for NAFLD.The NAFLD classification algorithm is superior to ICD-9 billing data alone. This approach is simple to develop, deploy, and can be applied across different institutions to create EMR-based cohorts of individuals with NAFLD.