Classifier performance prediction for computer-aided diagnosis using a limited dataset

Classifier performance prediction for computer-aided diagnosis using a limited dataset
复制标题

DOI:
10.1118/1.2868757
复制
发表时间:
2008-04-01
期刊:
影响因子:
3.8
通讯作者:
Hadjiiski, Lubomir
Hadjiiski, Lubomir
中科院分区:
医学3区
文献类型:
--
作者:
Sahiner, Berkman;Chan, Heang-Ping;Hadjiiski, Lubomir

文献摘要

被引文献

相似文献

在实际的分类器设计问题中,真实总体通常是未知的,并且可用样本的大小是有限的。一种常见的方法是使用重采样技术来估计将使用可用样本进行训练的分类器的性能。我们进行了蒙特卡罗模拟研究,以比较不同重采样技术在有限大小样本约束下训练分类器和预测其性能的能力。假设这两类的真实总体是具有已知协方差矩阵的多元正态分布。从总体中抽取有限的样本向量集。分类器的真实性能定义为当用特定样本设计的分类器应用于真实总体时,受试者工作特征曲线(AUC)下的面积。我们研究了基于 Fukunaga-Hayes 和留一法技术的方法,以及三种不同类型的引导方法,即普通引导方法、0.632 引导方法和 0.632+ 引导方法。使用费舍尔线性判别分析作为分类器。特征空间的维数从 3 到 15 不等。正类的样本大小 n(2) 在 25 到 60 之间变化,而负类的案例数量等于 n(2) 或 3n(2)。每个实验都是使用从真实人群中随机抽取的独立数据集进行的。我们对每个模拟条件总共进行了 1000 次实验,比较了使用不同重采样技术估计的 AUC 的偏差、方差和均方根误差 (RMSE) 与真实 AUC(通过有限数据集训练和总体测试获得)。我们的结果表明,在研究条件下,使用不同重采样方法获得的 RMSE 可能存在较大差异,特别是当特征空间维数较大且样本量较小时。在这种情况下,0.632 和 0.632+ bootstrap 方法具有最低的 RMSE,这表明使用 0.632 和 0.632+ bootstrap 获得的估计性能与真实性能之间的差异在统计上将小于使用其他三种重采样方法获得的差异。在三种引导方法中,0.632+引导方法提供最低的偏差。尽管这项研究是在某些特定条件下进行的,但它揭示了有限数据集约束下分类器性能预测问题的重要趋势。 (c) 2008 年美国医学物理学家协会。
In a practical classifier design problem, the true population is generally unknown and the available sample is finite-sized. A common approach is to use a resampling technique to estimate the performance of the classifier that will be trained with the available sample. We conducted a Monte Carlo simulation study to compare the ability of the different resampling techniques in training the classifier and predicting its performance under the constraint of a finite-sized sample. The true population for the two classes was assumed to be multivariate normal distributions with known covariance matrices. Finite sets of sample vectors were drawn from the population. The true performance of the classifier is defined as the area under the receiver operating characteristic curve (AUC) when the classifier designed with the specific sample is applied to the true population. We investigated methods based on the Fukunaga-Hayes and the leave-one-out techniques, as well as three different types of bootstrap methods, namely, the ordinary, 0.632, and 0.632+ bootstrap. The Fisher's linear discriminant analysis was used as the classifier. The dimensionality of the feature space was varied from 3 to 15. The sample size n(2) from the positive class was varied between 25 and 60, while the number of cases from the negative class was either equal to n(2) or 3n(2). Each experiment was performed with an independent dataset randomly drawn from the true population. Using a total of 1000 experiments for each simulation condition, we compared the bias, the variance, and the root-mean-squared error (RMSE) of the AUC estimated using the different resampling techniques relative to the true AUC (obtained from training on a finite dataset and testing on the population). Our results indicated that, under the study conditions, there can be a large difference in the RMSE obtained using different resampling methods, especially when the feature space dimensionality is relatively large and the sample size is small. Under this type of conditions, the 0.632 and 0.632+ bootstrap methods have the lowest RMSE, indicating that the difference between the estimated and the true performances obtained using the 0.632 and 0.632+ bootstrap will be statistically smaller than those obtained using the other three resampling methods. Of the three bootstrap methods, the 0.632+ bootstrap provides the lowest bias. Although this investigation is performed under some specific conditions, it reveals important trends for the problem of classifier performance prediction under the constraint of a limited dataset. (c) 2008 American Association of Physicists in Medicine.