Medical subdomain classification of clinical notes using a machine learning-based natural language processing approach.

Medical subdomain classification of clinical notes using a machine learning-based natural language processing approach.
复制标题

DOI:
10.1186/s12911-017-0556-8
复制
发表时间:
2017-12-01
影响因子:
3.5
通讯作者:
Chueh HC
Chueh HC
中科院分区:
医学3区
文献类型:
--
作者:
Weng WH;Wagholikar KB;McCray AT;Szolovits P;Chueh HC

文献摘要

参考文献

被引文献

相似文献

临床笔记的医学子域(如心脏病学或神经病学)是用于开发机器学习下游应用程序的有用内容派生元数据。为了准确地对笔记的医学子域进行分类,我们构建了一个基于机器学习的自然语言处理(NLP)管道,并根据笔记的内容开发了医学子域分类器。我们使用临床NLP系统、临床文本分析和知识提取系统(cTAKES)、统一医学语言系统(UMLS)元词库、语义网络和学习算法构建了管道,以从两个数据集-来自整合数据分析、分析、和共享(iDASH)数据存储库(n = 431)和马萨诸塞州总医院(MGH)(n = 91,237),并使用数据表示方法和监督学习算法的不同组合构建医学子域分类器。我们评估了分类器的性能及其在两个数据集上的可移植性。具有神经词嵌入训练医学子域分类器的卷积递归神经网络在iDASH和MGH数据集上产生了最佳性能测量,接收器工作特征曲线(AUC)下的面积分别为0.975和0.991,F1得分分别为0.845和0.870。考虑到更好的临床可解释性,线性支持向量机训练的医学子域分类器使用混合词袋和临床相关的UMLS概念作为特征表示,具有词频-逆文档频率(tf-idf)加权,在iDASH和MGH数据集上优于其他浅层学习分类器,AUC为0.957和0.964,F1评分分别为0.932和0.934。我们在一个数据集上训练分类器,应用于另一个数据集,并在我们研究的一半医学子域的分类器中产生了F1分数为0.7的阈值。我们的研究表明,基于监督学习的自然语言处理方法是有用的开发医学子域分类器。具有分布式单词表示的深度学习算法具有更好的性能,而具有单词和概念表示的浅层学习算法具有更好的临床可解释性。便携式分类器也可以跨来自不同机构的数据集使用。本文的在线版本(10.1186/s12911-017-0556-8)包含补充材料,可供授权用户使用。
The medical subdomain of a clinical note, such as cardiology or neurology, is useful content-derived metadata for developing machine learning downstream applications. To classify the medical subdomain of a note accurately, we have constructed a machine learning-based natural language processing (NLP) pipeline and developed medical subdomain classifiers based on the content of the note. We constructed the pipeline using the clinical NLP system, clinical Text Analysis and Knowledge Extraction System (cTAKES), the Unified Medical Language System (UMLS) Metathesaurus, Semantic Network, and learning algorithms to extract features from two datasets — clinical notes from Integrating Data for Analysis, Anonymization, and Sharing (iDASH) data repository (n = 431) and Massachusetts General Hospital (MGH) (n = 91,237), and built medical subdomain classifiers with different combinations of data representation methods and supervised learning algorithms. We evaluated the performance of classifiers and their portability across the two datasets. The convolutional recurrent neural network with neural word embeddings trained-medical subdomain classifier yielded the best performance measurement on iDASH and MGH datasets with area under receiver operating characteristic curve (AUC) of 0.975 and 0.991, and F1 scores of 0.845 and 0.870, respectively. Considering better clinical interpretability, linear support vector machine-trained medical subdomain classifier using hybrid bag-of-words and clinically relevant UMLS concepts as the feature representation, with term frequency-inverse document frequency (tf-idf)-weighting, outperformed other shallow learning classifiers on iDASH and MGH datasets with AUC of 0.957 and 0.964, and F1 scores of 0.932 and 0.934 respectively. We trained classifiers on one dataset, applied to the other dataset and yielded the threshold of F1 score of 0.7 in classifiers for half of the medical subdomains we studied. Our study shows that a supervised learning-based NLP approach is useful to develop medical subdomain classifiers. The deep learning algorithm with distributed word representation yields better performance yet shallow learning algorithms with the word and concept representation achieves comparable performance with better clinical interpretability. Portable classifiers may also be used across datasets from different institutions. The online version of this article (10.1186/s12911-017-0556-8) contains supplementary material, which is available to authorized users.
药物不良事件的文本挖掘:前景、挑战和最新技术。
DOI: 10.1007/s40264-014-0218-z
发表时间: 2014-10
期刊: DRUG SAFETY
影响因子: 4.2
作者:
Harpaz, Rave;Callahan, Alison;Tamang, Suzanne;Low, Yen;Odgers, David;Finlayson, Sam;Jung, Kenneth;LePendu, Paea;Shah, Nigam H.
通讯作者: Shah, Nigam H.
DOI: 10.1371/journal.pone.0136651
发表时间: 2015
期刊: PloS one
影响因子: 3.7
作者:
Liao KP;Ananthakrishnan AN;Kumar V;Xia Z;Cagan A;Gainer VS;Goryachev S;Chen P;Savova GK;Agniel D;Churchill S;Lee J;Murphy SN;Plenge RM;Szolovits P;Kohane I;Shaw SY;Karlson EW;Cai T
通讯作者: Cai T
DOI: 10.1371/journal.pone.0087555
发表时间: 2014
期刊: PloS one
影响因子: 3.7
作者:
Cohen R;Aviram I;Elhadad M;Elhadad N
通讯作者: Elhadad N
DOI: 10.1016/j.ijmedinf.2012.12.005
发表时间: 2014-12
影响因子: 4.9
作者:
Byrd RJ;Steinhubl SR;Sun J;Ebadollahi S;Stewart WF
通讯作者: Stewart WF
DOI: 10.3233/978-1-61499-753-5-246
发表时间: 2017-01-01
期刊: INFORMATICS FOR HEALTH: CONNECTED CITIZEN-LED WELLNESS AND POPULATION HEALTH
影响因子: --
作者:
Hughes, Mark;Li, Irene;Suzumura, Toyotaro
通讯作者: Suzumura, Toyotaro