Comparing deep learning and concept extraction based methods for patient phenotyping from clinical narratives.

Comparing deep learning and concept extraction based methods for patient phenotyping from clinical narratives.
复制标题

DOI:
10.1371/journal.pone.0192360
复制
发表时间:
2018
期刊:
影响因子:
3.7
通讯作者:
Celi LA
Celi LA
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Gehrmann S;Dernoncourt F;Li Y;Carlson ET;Wu JT;Welt J;Foote J Jr;Moseley ET;Grant DW;Tyler PD;Celi LA

文献摘要

参考文献

被引文献

相似文献

在电子健康记录的二次分析中,一项关键任务是正确识别正在调查的患者队列。在许多情况下,最有价值的和相关的信息,以准确分类的医疗条件只存在于临床叙述。因此,有必要使用自然语言处理(NLP)技术来提取和评估这些叙述。解决这个问题最常用的方法是从文本中提取一些临床医生定义的医学概念,并使用机器学习技术来识别特定患者是否患有某种疾病。然而,深度学习和NLP的最新进展使模型能够学习(医学)语言的丰富表示。用于文本分类的卷积神经网络(CNN)可以通过利用语言的表示来学习文本中的哪些短语与给定的医疗条件相关,从而增强现有技术。在这项工作中,我们使用MIMIC-III数据库中的1,610份出院摘要,在10个表型任务中比较了基于概念提取的方法与CNN和其他常用的NLP模型。我们发现,CNN在几乎所有的任务中都优于基于概念提取的方法,F1分数提高了26个百分点,ROC曲线下面积(AUC)提高了7个百分点。我们还评估了这两种方法的可解释性,提出和评估的方法,计算和提取最突出的短语的预测。结果表明,CNN是患者表型和队列识别中现有方法的有效替代方案,应该进一步研究。此外,本文中提出的深度学习方法可用于在图表审查期间帮助临床医生,或通过识别和突出显示各种医疗条件的相关短语来支持从文本中提取账单代码。
In secondary analysis of electronic health records, a crucial task consists in correctly identifying the patient cohort under investigation. In many cases, the most valuable and relevant information for an accurate classification of medical conditions exist only in clinical narratives. Therefore, it is necessary to use natural language processing (NLP) techniques to extract and evaluate these narratives. The most commonly used approach to this problem relies on extracting a number of clinician-defined medical concepts from text and using machine learning techniques to identify whether a particular patient has a certain condition. However, recent advances in deep learning and NLP enable models to learn a rich representation of (medical) language. Convolutional neural networks (CNN) for text classification can augment the existing techniques by leveraging the representation of language to learn which phrases in a text are relevant for a given medical condition. In this work, we compare concept extraction based methods with CNNs and other commonly used models in NLP in ten phenotyping tasks using 1,610 discharge summaries from the MIMIC-III database. We show that CNNs outperform concept extraction based methods in almost all of the tasks, with an improvement in F1-score of up to 26 and up to 7 percentage points in area under the ROC curve (AUC). We additionally assess the interpretability of both approaches by presenting and evaluating methods that calculate and extract the most salient phrases for a prediction. The results indicate that CNNs are a valid alternative to existing approaches in patient phenotyping and cohort identification, and should be further investigated. Moreover, the deep learning approach presented in this paper can be used to assist clinicians during chart review or support the extraction of billing codes from text by identifying and highlighting relevant phrases for various medical conditions.
DOI: 10.1177/0272989x11400418
发表时间: 2012-01
影响因子: 3.6
作者:
Denny, Joshua C.;Choma, Neesha N.;Peterson, Josh F.;Miller, Randolph A.;Bastarache, Lisa;Li, Ming;Peterson, Neeraja B.
通讯作者: Peterson, Neeraja B.
DOI: 10.1016/j.mayocp.2016.08.008
发表时间: 2016-11-01
影响因子: 8.9
作者:
Ackerman, Jaeger P.;Bartos, Daniel C.;Ackerman, Michael J.
通讯作者: Ackerman, Michael J.
DOI: 10.1093/jamia/ocv155
发表时间: 2016-04-01
影响因子: 6.4
作者:
Bates, Jonathan;Fodeh, Samah J.;Womack, Julie A.
通讯作者: Womack, Julie A.
DOI: 10.1097/mib.0b013e31828133fd
发表时间: 2013-06
影响因子: 4.9
作者:
Ananthakrishnan AN;Cai T;Savova G;Cheng SC;Chen P;Perez RG;Gainer VS;Murphy SN;Szolovits P;Xia Z;Shaw S;Churchill S;Karlson EW;Kohane I;Plenge RM;Liao KP
通讯作者: Liao KP
DOI: 10.1093/jamia/ocw156
发表时间: 2017-05-01
影响因子: 6.4
作者:
Dernoncourt, Franck;Lee, Ji Young;Szolovits, Peter
通讯作者: Szolovits, Peter