Risks of feature leakage and sample size dependencies in deep feature extraction for breast mass classification.
Risks of feature leakage and sample size dependencies in deep feature extraction for breast mass classification.
复制标题
DOI:
10.1002/mp.14678
复制
发表时间:
2021-06
期刊:
影响因子:
3.8
通讯作者:
Helvie MA
中科院分区:
文献类型:
--
作者:
Samala RK;Chan HP;Hadjiiski L;Helvie MA
Transfer learning is commonly used in deep learning for medical imaging to alleviate the problem of limited available data. In this work we studied the risk of feature leakage and its dependence on sample size when using pre-trained deep convolutional neural network (DCNN) as feature extractor for classification breast masses in mammography. Feature leakage occurs when the training set is used for feature selection and classifier modeling while the cost function is guided by the validation performance or informed by the test performance. The high-dimensional feature space extracted from pre-trained DCNN suffers from the curse of dimensionality; feature subsets that can provide excessively optimistic performance can be found for the validation set or test set if the latter is allowed for unlimited reuse during algorithm development. We designed a simulation study to examine feature leakage when using DCNN as feature extractor for mass classification in mammography. 4,577 unique mass lesions were partitioned by patient into three sets: 3,222 for training, 508 for validation and 847 for independent testing. Three pre-trained DCNNs, AlexNet, GoogLeNet, and VGG16, were first compared using a training set in four-fold cross validation and one was selected as the feature extractor. To assess generalization errors, the independent test set was sequestered as truly unseen cases. A training set of a range of sizes from 10% to 75% was simulated by random drawing from the available training set in addition to 100% of the training set. Three commonly used feature classifiers, the linear discriminant, the support vector machine, and the random forest were evaluated. A sequential feature selection method was used to find feature subsets that could achieve high classification performance in terms of the area under the receiver operating characteristic curve (AUC) in the validation set. The extent of feature leakage and the impact of training set size were analyzed by comparison to the performance in the unseen test set. All three classifiers showed large generalization error between the validation set and the independent sequestered test set at all sample sizes. The generalization error decreased as the sample size increased. At 100% of the sample size, one classifier achieved an AUC as high as 0.91 on the validation set while the corresponding performance on the unseen test set only reached an AUC of 0.72. Our results demonstrate that large generalization errors can occur in AI tools due to feature leakage. Without evaluation on unseen test cases, optimistically biased performance may be reported inadvertently, and can lead to unrealistic expectations and reduce confidence for clinical implementation.
登录
查看更多内容
影响因子:
3.5
作者:
Samala RK;Chan HP;Hadjiiski LM;Helvie MA;Cha KH;Richter CD
通讯作者:
Richter CD
影响因子:
3.8
作者:
Samala, Ravi K.;Chan, Heang-Ping;Cha, Kenny
通讯作者:
Cha, Kenny
影响因子:
9.8
作者:
Lee RS;Gimenez F;Hoogi A;Miyake KK;Gorovoy M;Rubin DL
通讯作者:
Rubin DL
影响因子:
4.6
作者:
Lao J;Chen Y;Li ZC;Li Q;Zhang J;Liu J;Zhai G
通讯作者:
Zhai G
影响因子:
5.8
作者:
Molinaro, AM;Simon, R;Pfeiffer, RM
通讯作者:
Pfeiffer, RM