Improving breast mass classification by shared data with domain transformation using a generative adversarial network

Improving breast mass classification by shared data with domain transformation using a generative adversarial network
复制标题

DOI:
10.1016/j.compbiomed.2020.103698
复制
发表时间:
2020-04-01
影响因子:
7.7
通讯作者:
Fujita, Hiroshi
Fujita, Hiroshi
中科院分区:
工程技术2区
文献类型:
--
作者:
Muramatsu, Chisako;Nishio, Mizuho;Fujita, Hiroshi

文献摘要

被引文献

相似文献

卷积神经网络(CNN)的训练通常需要大量的数据集。然而,收集大型医学图像数据集并不容易。本研究的目的是研究合成图像在训练CNN中的实用性,并通过域变换证明不相关图像的适用性。乳房X线片显示202良性和212恶性肿块用于评价。为了创建合成数据,使用计算机断层扫描(CT)中的599个肺结节和数字化乳房X线照片(DDSM)中的1430个乳房肿块训练了一个循环生成对抗网络。CNN被训练用于良性和恶性肿块之间的分类。在使用原始数据、增强数据、合成数据、DDSM图像和自然图像(ImageNet数据集)训练的网络之间比较分类性能。根据分类准确性和受试者工作特征曲线下面积(AUC)对结果进行评价。随着数据的增加,分类准确率从65.7%提高到67.1%。使用ImageNet预训练模型是有用的(79.2%)。当合成图像或DDSM图像仅用于预训练时,性能略有改善(分别为67.6%和72.5%)。当使用合成图像训练ImageNet预训练模型时,分类性能略有提高(81.4%),尽管AUC的差异在统计学上并不显著。合成图像的使用具有与DDSM图像类似的效果。研究结果表明,通过域变换从无关病变生成的合成数据可以用于增加训练样本。
Training of a convolutional neural network (CNN) generally requires a large dataset. However, it is not easy to collect a large medical image dataset. The purpose of this study is to investigate the utility of synthetic images in training CNNs and to demonstrate the applicability of unrelated images by domain transformation. Mammograms showing 202 benign and 212 malignant masses were used for evaluation. To create synthetic data, a cycle generative adversarial network was trained with 599 lung nodules in computed tomography (CT) and 1430 breast masses on digitized mammograms (DDSM). A CNN was trained for classification between benign and malignant masses. The classification performance was compared between the networks trained with the original data, augmented data, synthetic data, DDSM images, and natural images (ImageNet dataset). The results were evaluated in terms of the classification accuracy and the area under the receiver operating characteristic curves (AUC). The classification accuracy improved from 65.7% to 67.1% with data augmentation. The use of an ImageNet pretrained model was useful (79.2%). Performance was slightly improved when synthetic images or the DDSM images only were used for pretraining (67.6 and 72.5%, respectively). When the ImageNet pretrained model was trained with the synthetic images, the classification performance slightly improved (81.4%), although the difference in AUCs was not statistically significant. The use of the synthetic images had an effect similar to the DDSM images. The results of the proposed study indicated that the synthetic data generated from unrelated lesions by domain transformation could be used to increase the training samples.