Risks of feature leakage and sample size dependencies in deep feature extraction for breast mass classification.

Risks of feature leakage and sample size dependencies in deep feature extraction for breast mass classification.
复制标题

DOI:
10.1002/mp.14678
复制
发表时间:
2021-06
期刊:
影响因子:
3.8
通讯作者:
Helvie MA
Helvie MA
中科院分区:
医学3区
文献类型:
--
作者:
Samala RK;Chan HP;Hadjiiski L;Helvie MA

文献摘要

参考文献

被引文献

相似文献

迁移学习通常用于医学成像的深度学习,以缓解可用数据有限的问题。在这项工作中,我们研究了使用预训练的深度卷积神经网络(DCNN)作为乳房X光检查中乳腺肿块分类的特征提取器时特征泄漏的风险及其对样本量的依赖性。当训练集用于特征选择和分类器建模时,会发生特征泄漏,而成本函数由验证性能指导或由测试性能通知。从预训练的DCNN中提取的高维特征空间受到维数灾难的影响;可以为验证集或测试集找到可以提供过度乐观性能的特征子集,如果后者在算法开发期间允许无限重用。我们设计了一项模拟研究,以检查使用DCNN作为乳腺X射线摄影中肿块分类的特征提取器时的特征泄漏。将4,577个独特的肿块病变按患者分为三组:3,222个用于训练,508个用于验证,847个用于独立测试。三个预训练的DCNN,AlexNet,GoogLeNet和VGG 16,首先使用四重交叉验证的训练集进行比较,并选择一个作为特征提取器。为了评估泛化错误,独立测试集被隔离为真正看不见的案例。除了100%的训练集之外,还通过从可用的训练集随机抽取来模拟10%至75%的大小范围的训练集。三种常用的特征分类器,线性判别,支持向量机,和随机森林进行了评估。一个顺序的特征选择方法被用来找到功能子集,可以实现高的分类性能的验证集的受试者工作特征曲线(AUC)下的面积。特征泄漏的程度和训练集大小的影响通过与看不见的测试集中的性能进行比较来分析。所有三个分类器在所有样本量下的验证集和独立隔离测试集之间都存在较大的泛化误差。泛化误差随着样本量的增加而减小。在100%的样本量下,一个分类器在验证集上实现了高达0.91的AUC,而在看不见的测试集上的相应性能仅达到0.72的AUC。我们的研究结果表明,由于特征泄漏,AI工具可能会出现较大的泛化错误。如果不对看不见的测试用例进行评估,可能会无意中报告乐观偏倚的性能,并可能导致不切实际的期望,降低临床实施的信心。
Transfer learning is commonly used in deep learning for medical imaging to alleviate the problem of limited available data. In this work we studied the risk of feature leakage and its dependence on sample size when using pre-trained deep convolutional neural network (DCNN) as feature extractor for classification breast masses in mammography. Feature leakage occurs when the training set is used for feature selection and classifier modeling while the cost function is guided by the validation performance or informed by the test performance. The high-dimensional feature space extracted from pre-trained DCNN suffers from the curse of dimensionality; feature subsets that can provide excessively optimistic performance can be found for the validation set or test set if the latter is allowed for unlimited reuse during algorithm development. We designed a simulation study to examine feature leakage when using DCNN as feature extractor for mass classification in mammography. 4,577 unique mass lesions were partitioned by patient into three sets: 3,222 for training, 508 for validation and 847 for independent testing. Three pre-trained DCNNs, AlexNet, GoogLeNet, and VGG16, were first compared using a training set in four-fold cross validation and one was selected as the feature extractor. To assess generalization errors, the independent test set was sequestered as truly unseen cases. A training set of a range of sizes from 10% to 75% was simulated by random drawing from the available training set in addition to 100% of the training set. Three commonly used feature classifiers, the linear discriminant, the support vector machine, and the random forest were evaluated. A sequential feature selection method was used to find feature subsets that could achieve high classification performance in terms of the area under the receiver operating characteristic curve (AUC) in the validation set. The extent of feature leakage and the impact of training set size were analyzed by comparison to the performance in the unseen test set. All three classifiers showed large generalization error between the validation set and the independent sequestered test set at all sample sizes. The generalization error decreased as the sample size increased. At 100% of the sample size, one classifier achieved an AUC as high as 0.91 on the validation set while the corresponding performance on the unseen test set only reached an AUC of 0.72. Our results demonstrate that large generalization errors can occur in AI tools due to feature leakage. Without evaluation on unseen test cases, optimistically biased performance may be reported inadvertently, and can lead to unrealistic expectations and reduce confidence for clinical implementation.
DOI: 10.1088/1361-6560/aa93d4
发表时间: 2017-11-10
影响因子: 3.5
作者:
Samala RK;Chan HP;Hadjiiski LM;Helvie MA;Cha KH;Richter CD
通讯作者: Richter CD
DOI: 10.1118/1.4967345
发表时间: 2016-12-01
期刊: MEDICAL PHYSICS
影响因子: 3.8
作者:
Samala, Ravi K.;Chan, Heang-Ping;Cha, Kenny
通讯作者: Cha, Kenny
DOI: 10.1038/sdata.2017.177
发表时间: 2017-12-19
期刊: Scientific data
影响因子: 9.8
作者:
Lee RS;Gimenez F;Hoogi A;Miyake KK;Gorovoy M;Rubin DL
通讯作者: Rubin DL
基于深度学习的放射组学模型,用于预测多形性胶质母细胞瘤的生存期
DOI: 10.1038/s41598-017-10649-8
发表时间: 2017-09-04
期刊: Scientific reports
影响因子: 4.6
作者:
Lao J;Chen Y;Li ZC;Li Q;Zhang J;Liu J;Zhai G
通讯作者: Zhai G
DOI: 10.1093/bioinformatics/bti499
发表时间: 2005-08-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Molinaro, AM;Simon, R;Pfeiffer, RM
通讯作者: Pfeiffer, RM