Convolutional Neural Networks for Medical Image Analysis: Full Training or Fine Tuning?

Convolutional Neural Networks for Medical Image Analysis: Full Training or Fine Tuning?
复制标题

DOI:
10.1109/tmi.2016.2535302
复制
发表时间:
2016-05-01
影响因子:
10.6
通讯作者:
Liang, Jianming
Liang, Jianming
中科院分区:
工程技术1区
文献类型:
--
作者:
Tajbakhsh, Nima;Shin, Jae Y.;Liang, Jianming

文献摘要

被引文献

相似文献

从头开始训练深度卷积神经网络(CNN)是很困难的,因为它需要大量的标记训练数据和大量的专业知识来确保正确的收敛。一个有希望的替代方案是微调已经使用例如大量标记的自然图像进行预训练的CNN。然而,自然图像和医学图像之间的实质性差异可能会反对这种知识转移。在本文中,我们试图在医学图像分析的背景下回答以下中心问题:使用经过充分微调的预训练深度CNN是否可以消除从头开始训练深度CNN的需要?为了解决这个问题,我们考虑了三个专业(放射学,心脏病学和胃肠病学)中的四种不同的医学成像应用,涉及三种不同成像模式的分类,检测和分割,并研究了从头开始训练的深度CNN的性能如何与以分层方式微调的预训练CNN进行比较。我们的实验一致表明:1)使用经过充分微调的预训练CNN优于从头开始训练的CNN,或者在最坏的情况下,表现得与从头开始训练的CNN一样好; 2)微调CNN比从头开始训练的CNN对训练集的大小更鲁棒; 3)浅调和深调都不是特定应用的最佳选择;以及4)我们的逐层微调方案可以提供一种实用的方法,以根据可用数据的量来达到当前应用的最佳性能。
Training a deep convolutional neural network (CNN) from scratch is difficult because it requires a large amount of labeled training data and a great deal of expertise to ensure proper convergence. A promising alternative is to fine-tune a CNN that has been pre-trained using, for instance, a large set of labeled natural images. However, the substantial differences between natural and medical images may advise against such knowledge transfer. In this paper, we seek to answer the following central question in the context of medical image analysis: Can the use of pre-trained deep CNNs with sufficient fine-tuning eliminate the need for training a deep CNN from scratch? To address this question, we considered four distinct medical imaging applications in three specialties (radiology, cardiology, and gastroenterology) involving classification, detection, and segmentation from three different imaging modalities, and investigated how the performance of deep CNNs trained from scratch compared with the pre-trained CNNs fine-tuned in a layer-wise manner. Our experiments consistently demonstrated that 1) the use of a pre-trained CNN with adequate fine-tuning outperformed or, in the worst case, performed as well as a CNN trained from scratch; 2) fine-tuned CNNs were more robust to the size of training sets than CNNs trained from scratch; 3) neither shallow tuning nor deep tuning was the optimal choice for a particular application; and 4) our layer-wise fine-tuning scheme could offer a practical way to reach the best performance for the application at hand based on the amount of available data.