Natural Language-based Machine Learning Models for the Annotation of Clinical Radiology Reports

Natural Language-based Machine Learning Models for the Annotation of Clinical Radiology Reports
复制标题

DOI:
10.1148/radiol.2018171093
复制
发表时间:
2018-05-01
期刊:
影响因子:
19.7
通讯作者:
Oermann, Eric Karl
Oermann, Eric Karl
中科院分区:
医学1区
文献类型:
--
作者:
Zech, John;Pain, Margaret;Oermann, Eric Karl

文献摘要

被引文献

相似文献

目的:比较不同的方法产生的放射学报告的功能,并开发一种方法来自动识别这些reports.Materials和方法的结果:在这项研究中,96,303头部计算机断层扫描(CT)的报告。这些报告的语言复杂性进行了比较,与替代语料库。头部CT报告进行预处理,并通过使用词袋(BOW),词嵌入,和潜在的Dirichlet分配为基础的方法构建机器可分析的功能。最终,1004份头部CT报告被医生手动标记为感兴趣的发现,其中一部分被认为是关键发现。Lasso逻辑回归用于训练模型,用于使用构建的特征对1004份头部CT报告中的602份(60%)进行医生分配的标签,并在1004份报告中的402份(40%)上验证这些模型的性能。通过受试者工作特征曲线下面积(AUC)对模型进行评分,并报告报告(a)所有标签、(B)关键标签和(c)报告中是否存在任何关键发现的合计AUC统计量。灵敏度、特异性、准确性和F1得分被报告用于最佳性能模型的(a)所有标签的预测和(B)包含关键发现的报告的识别。表现最好的模型(BOW与一元,二元,”(《说文解字》),字。用于识别任何关键头部CT发现的存在的AUC为0.966,所有头部CT发现的平均AUC为0.957。识别任何关键发现的灵敏度和特异性分别为92.59%(175/189)和89.67%(191/213)。所有结果的平均灵敏度和特异性分别为90.25%(1898/2103)和91.72%(18351/20007)。更简单的BOW方法取得的结果与更复杂的方法相比具有竞争力,对于一元模型BOW,任何关键发现的平均AUC为0.951,而对于性能最好的模型为0.966。头部CT语料库的Yule I为34,明显低于路透社语料库(103)或I2 B2出院摘要(271),表明较低的语言复杂性。这种方法的成功得益于这些报告的标准化语言。通过这种方法,可以为深度学习等应用生成大型标记语料库。(C)RSNA,2018
Purpose: To compare different methods for generating features from radiology reports and to develop a method to automatically identify findings in these reports.Materials and Methods: In this study, 96,303 head computed tomography (CT) reports were obtained. The linguistic complexity of these reports was compared with that of alternative corpora. Head CT reports were preprocessed, and machine-analyzable features were constructed by using bag-of-words (BOW), word embedding, and Latent Dirichlet allocation-based approaches. Ultimately, 1004 head CT reports were manually labeled for findings of interest by physicians, and a subset of these were deemed critical findings. Lasso logistic regression was used to train models for physician-assigned labels on 602 of 1004 head CT reports (60%) using the constructed features, and the performance of these models was validated on a held-out 402 of 1004 reports (40%). Models were scored by area under the receiver operating characteristic curve (AUC), and aggregate AUC statistics were reported for (a) all labels, (b) critical labels, and (c) the presence of any critical finding in a report. Sensitivity, specificity, accuracy, and F1 score were reported for the best performing model's (a) predictions of all labels and (b) identification of reports containing critical findings.Results: The best-performing model (BOW with unigrams, bigrams, and trigrams plus average word embeddings vector) had a held-out AUC of 0.966 for identifying the presence of any critical head CT finding and an average 0.957 AUC across all head CT findings. Sensitivity and specificity for identifying the presence of any critical finding were 92.59% (175 of 189) and 89.67% (191 of 213), respectively. Average sensitivity and specificity across all findings were 90.25% (1898 of 2103) and 91.72% (18 351 of 20 007), respectively. Simpler BOW methods achieved results competitive with those of more sophisticated approaches, with an average AUC for presence of any critical finding of 0.951 for unigram BOW versus 0.966 for the best-performing model. The Yule I of the head CT corpus was 34, markedly lower than that of the Reuters corpus (at 103) or I2B2 discharge summaries (at 271), indicating lower linguistic complexity.Conclusion: Automated methods can be used to identify findings in radiology reports. The success of this approach benefits from the standardized language of these reports. With this method, a large labeled corpus can be generated for applications such as deep learning. (C) RSNA, 2018