Predicting Classifier Performance with Limited Training Data: Applications to Computer-Aided Diagnosis in Breast and Prostate Cancer

Predicting Classifier Performance with Limited Training Data: Applications to Computer-Aided Diagnosis in Breast and Prostate Cancer
复制标题

DOI:
10.1371/journal.pone.0117900
复制
发表时间:
2015-05-18
期刊:
影响因子:
3.7
通讯作者:
Madabhushi, Anant
Madabhushi, Anant
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Basavanhally, Ajay;Viswanath, Satish;Madabhushi, Anant

文献摘要

被引文献

相似文献

临床试验越来越多地将医学成像数据与监督分类器结合使用,后者需要大量的训练数据来准确地对系统进行建模。然而,在试验开始时基于更小且更易于访问的数据集选择的分类器可能会产生不准确且不稳定的分类性能。在本文中,我们的目标是解决临床试验分类器选择中的两个常见问题:(1)根据较小数据集计算的错误率预测大型数据集的预期分类器性能;(2)根据较大数据集的预期性能选择适当的分类器。我们提出了一个框架,通过使用随机重复采样(RRS)与交叉验证采样策略相结合,仅使用有限数量的训练数据来对分类器进行比较评估。随后通过与在较大数据集上执行的留一交叉验证进行比较来验证推断的错误率。随着数据集大小的增加,预测错误率的能力在合成数据以及三种不同的计算成像任务上得到了证明:检测前列腺组织病理学中的癌性图像区域、区分乳腺组织病理学中的高级别和低级别癌症以及检测前列腺磁共振波谱中的癌性元体素。对于每个任务,都会探索 3 个不同分类器(k 最近邻、朴素贝叶斯、支持向量机)之间的关系。根据四分位距 (IQR) 进行的进一步定量评估表明,与不对所有三个数据集采用交叉验证抽样的传统 RRS 方法(平均 IQR 为 0.0297、0.0779 和 0.305)相比,我们的方法始终产生具有较低变异性的错误率(平均 IQR 为 0.0070、0.0127 和 0.0140)。
Clinical trials increasingly employ medical imaging data in conjunction with supervised classifiers, where the latter require large amounts of training data to accurately model the system. Yet, a classifier selected at the start of the trial based on smaller and more accessible datasets may yield inaccurate and unstable classification performance. In this paper, we aim to address two common concerns in classifier selection for clinical trials: (1) predicting expected classifier performance for large datasets based on error rates calculated from smaller datasets and (2) the selection of appropriate classifiers based on expected performance for larger datasets. We present a framework for comparative evaluation of classifiers using only limited amounts of training data by using random repeated sampling (RRS) in conjunction with a cross-validation sampling strategy. Extrapolated error rates are subsequently validated via comparison with leave-one-out cross-validation performed on a larger dataset. The ability to predict error rates as dataset size increases is demonstrated on both synthetic data as well as three different computational imaging tasks: detecting cancerous image regions in prostate histopathology, differentiating high and low grade cancer in breast histopathology, and detecting cancerous metavoxels in prostate magnetic resonance spectroscopy. For each task, the relationships between 3 distinct classifiers (k-nearest neighbor, naive Bayes, Support Vector Machine) are explored. Further quantitative evaluation in terms of interquartile range (IQR) suggests that our approach consistently yields error rates with lower variability (mean IQRs of 0.0070, 0.0127, and 0.0140) than a traditional RRS approach (mean IQRs of 0.0297, 0.0779, and 0.305) that does not employ cross-validation sampling for all three datasets.