Improving case definition of Crohn's disease and ulcerative colitis in electronic medical records using natural language processing: a novel informatics approach.

Improving case definition of Crohn's disease and ulcerative colitis in electronic medical records using natural language processing: a novel informatics approach.
复制标题

DOI:
10.1097/mib.0b013e31828133fd
复制
发表时间:
2013-06
影响因子:
4.9
通讯作者:
Liao KP
Liao KP
中科院分区:
医学2区
文献类型:
--
作者:
Ananthakrishnan AN;Cai T;Savova G;Cheng SC;Chen P;Perez RG;Gainer VS;Murphy SN;Szolovits P;Xia Z;Shaw S;Churchill S;Karlson EW;Kohane I;Plenge RM;Liao KP

文献摘要

被引文献

相似文献

以前的研究确定患者炎症性肠病(IBD)利用行政代码产生了不一致的结果。我们的目标是开发一个强大的基于电子病历(EMR)的IBD分类模型,利用编码数据和使用自然语言处理(NLP)的临床文本注释信息的组合。使用2个大型学术中心的EMR,我们创建了克罗恩病(CD)和溃疡性结肠炎(UC)的数据集市,包括每种疾病具有≥ 1个ICD-9编码的患者。我们利用来自临床记录的编码(即ICD 9代码、电子处方)和叙述性数据来开发我们的分类模型。模型开发和验证是在每种疾病随机选择的600名患者的训练集中进行的,医疗记录审查作为金标准。Logistic回归与自适应LASSO惩罚被用来选择信息变量。我们在CD训练集中确认了399例(67%)CD病例,在UC训练集中确认了378例(63%)UC病例。对于两者,包括叙述和编码数据的组合模型的准确性(CD的曲线下面积(AUC)0.95; UC 0.94)优于仅使用疾病ICD-9代码的模型(CD的AUC 0.89; UC的AUC 0.86)。在我们的最终模型中添加NLP叙述性术语,导致在相同准确度下多分类6-12%的受试者。与仅使用编码数据的模型相比,纳入使用NLP识别的叙事概念提高了CD和UC的EMR病例定义的准确性,同时识别了更多的受试者。
Prior studies identifying patients with inflammatory bowel disease (IBD) utilizing administrative codes have yielded inconsistent results. Our objective was to develop a robust electronic medical record (EMR) based model for classification of IBD leveraging the combination of codified data and information from clinical text notes using natural language processing (NLP). Using the EMR of 2 large academic centers, we created data marts for Crohn’s disease (CD) and ulcerative colitis (UC) comprising patients with ≥ 1 ICD-9 code for each disease. We utilized codified (i.e. ICD9 codes, electronic prescriptions) and narrative data from clinical notes to develop our classification model. Model development and validation was performed in a training set of 600 randomly selected patients for each disease with medical record review as the gold standard. Logistic regression with the adaptive LASSO penalty was used to select informative variables. We confirmed 399 (67%) CD cases in the CD training set and 378 (63%) UC cases in the UC training set. For both, a combined model including narrative and codified data had better accuracy (area under the curve (AUC) for CD 0.95; UC 0.94) than models utilizing only disease ICD-9 codes (AUC 0.89 for CD; 0.86 for UC). Addition of NLP narrative terms to our final model resulted in classification of 6–12% more subjects with the same accuracy. Inclusion of narrative concepts identified using NLP improves the accuracy of EMR case-definition for CD and UC while simultaneously identifying more subjects compared to models using codified data alone.