Comparative effectiveness of convolutional neural network (CNN) and recurrent neural network (RNN) architectures for radiology text report classification

Comparative effectiveness of convolutional neural network (CNN) and recurrent neural network (RNN) architectures for radiology text report classification
复制标题

卷积神经网络(CNN)和递归神经网络(RNN)结构在放射学文本报告分类中的有效性比较

DOI:
10.1016/j.artmed.2018.11.004
复制
发表时间:
2019-06-01
影响因子:
7.5
通讯作者:
Lungren, Matthew P.
Lungren, Matthew P.
中科院分区:
工程技术1区
文献类型:
--
作者:
Banerjee, Imon;Ling, Yuan;Lungren, Matthew P.

文献摘要

被引文献

相似文献

本文探讨了多机构尺度下医学影像自由文本报告信息提取的深度学习方法,并将它们与目前最先进的基于领域特定规则的系统PEFinder和传统的机器学习方法--支持向量机和Adaboost进行了比较。我们提出了两个不同的深度学习模型-(I)CNN Word-Glove和(Ii)基于领域短语注意力的分层递归神经网络(DPA-HNN),用于从四个主要医疗中心收集的7370多份临床胸部CT(CT)自由文本放射学报告中合成有关肺血栓(PE)的信息。我们提出的DPA-HNN模型将与领域相关的短语编码成一种注意机制,并通过由词级、句子级和文档级表示组成的分层RNN结构来表示放射学报告。实验结果表明,在单个机构数据集上训练的深度学习模型在我们的多机构测试集上的性能优于基于规则的PEFinder。在成人患者群体中存在PE的最佳F1得分为0.99(DPA-HNN),对于儿科人群的最佳F1得分为0.99(HNN),这表明基于成人数据训练的深度学习模型对儿科人群具有相当的准确性。我们的工作表明,在多机构成像文本报告的自动分类中广泛使用神经网络模型的可行性,用于各种应用,包括成像利用率评估、成像产量、临床决策支持工具,以及作为医学成像深度学习工作的大型语料库自动分类的一部分。
This paper explores cutting-edge deep learning methods for information extraction from medical imaging free text reports at a multi-institutional scale and compares them to the state-of-the-art domain-specific rule-based system - PEFinder and traditional machine learning methods - SVM and Adaboost. We proposed two distinct deep learning models - (i) CNN Word - Glove, and (ii) Domain phrase attention-based hierarchical recurrent neural network (DPA-HNN), for synthesizing information on pulmonary emboli (PE) from over 7370 clinical thoracic computed tomography (CT) free-text radiology reports collected from four major healthcare centers. Our proposed DPA-HNN model encodes domain-dependent phrases into an attention mechanism and represents a radiology report through a hierarchical RNN structure composed of word-level, sentence-level and document level representations. Experimental results suggest that the performance of the deep learning models that are trained on a single institutional dataset, are better than rule-based PEFinder on our multi-institutional test sets. The best F1 score for the presence of PE in an adult patient population was 0.99 (DPA-HNN) and for a pediatrics population was 0.99 (HNN) which shows that the deep learning models being trained on adult data, demonstrated generalizability to pediatrics population with comparable accuracy. Our work suggests feasibility of broader usage of neural network models in automated classification of multi-institutional imaging text reports for a variety of applications including evaluation of imaging utilization, imaging yield, clinical decision support tools, and as part of automated classification of large corpus for medical imaging deep learning work.