Machine-learning-based multiple abnormality prediction with large-scale chest computed tomography volumes.

Machine-learning-based multiple abnormality prediction with large-scale chest computed tomography volumes.
复制标题

DOI:
10.1016/j.media.2020.101857
复制
发表时间:
2021-01
影响因子:
10.9
通讯作者:
Carin L
Carin L
中科院分区:
工程技术1区
文献类型:
--
作者:
Draelos RL;Dov D;Mazurowski MA;Lo JY;Henao R;Rubin GD;Carin L

文献摘要

参考文献

被引文献

相似文献

放射学的机器学习模型受益于具有高质量异常标签的大规模数据集。我们整理和分析了来自19993名独特患者的36316卷的胸部计算机断层扫描(CT)数据集。这是已报道的最大的多注解体积医学成像数据集。为了标注这个数据集,我们开发了一种基于规则的方法,用于从自由文本放射学报告中自动提取异常标签,平均F-Score为0.976(最小0.941,最大1.0)。我们还开发了一个使用深度卷积神经网络(CNN)对胸部CT体积进行多器官、多疾病分类的模型。该模型对18种异常的分类性能达到0.90,对所有83种异常的平均AUROC为0.773,证明了从未过滤的整体CT数据中学习的可行性。我们表明,对更多标签的训练显著提高了性能:对于9个标签的子集-结节、不透明、肺不张、胸腔积液、实变、肿块、心包积液、心脏肿大和气胸-当训练标签的数量从9个增加到全部83个时,模型的平均AUROC增加了10%。所有用于体积预处理、自动标签提取和体积异常预测模型的代码都是公开可用的。36,316个CT卷和标签也将在机构批准之前公开提供。
Machine learning models for radiology benefit from large-scale data sets with high quality labels for abnormalities. We curated and analyzed a chest computed tomography (CT) data set of 36,316 volumes from 19,993 unique patients. This is the largest multiply-annotated volumetric medical imaging data set reported. To annotate this data set, we developed a rule-based method for automatically extracting abnormality labels from free-text radiology reports with an average F-score of 0.976 (min 0.941, max 1.0). We also developed a model for multi-organ, multi-disease classification of chest CT volumes that uses a deep convolutional neural network (CNN). This model reached a classification performance of AUROC > 0.90 for 18 abnormalities, with an average AUROC of 0.773 for all 83 abnormalities, demonstrating the feasibility of learning from unfiltered whole volume CT data. We show that training on more labels improves performance significantly: for a subset of 9 labels – nodule, opacity, atelectasis, pleural effusion, consolidation, mass, pericardial effusion, cardiomegaly, and pneumothorax – the model’s average AUROC increased by 10% when the number of training labels was increased from 9 to all 83. All code for volume preprocessing, automated label extraction, and the volume abnormality prediction model is publicly available. The 36,316 CT volumes and labels will also be made publicly available pending institutional approval.
DOI: 10.1148/radiol.09091308
发表时间: 2010-05-01
期刊: RADIOLOGY
影响因子: 19.7
作者:
de Hoop, Bartjan;Schaefer-Prokop, Cornelia;Prokop, Mathias
通讯作者: Prokop, Mathias
DOI: 10.1109/jbhi.2016.2636929
发表时间: 2017-01-01
影响因子: 7.7
作者:
Christodoulidis, Stergios;Anthimopoulos, Marios;Mougiakakou, Stavroula
通讯作者: Mougiakakou, Stavroula
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y
DOI: 10.1148/radiol.2018180763
发表时间: 2018-12-01
期刊: RADIOLOGY
影响因子: 19.7
作者:
Choi, Kyu Jin;Jang, Jong Keon;Yu, Eun Sil
通讯作者: Yu, Eun Sil
DOI: 10.1097/rct.0000000000000438
发表时间: 2016-09-01
影响因子: 1.3
作者:
Edwards, Rachael M.;Godwin, J. David;Kicska, Gregory
通讯作者: Kicska, Gregory