Bias in error estimation when using cross-validation for model selection.

Bias in error estimation when using cross-validation for model selection.
复制标题

DOI:
10.1186/1471-2105-7-91
复制
发表时间:
2006-02-23
期刊:
影响因子:
3
通讯作者:
Simon, R
Simon, R
中科院分区:
生物学4区
文献类型:
--
作者:
Varma, S;Simon, R

文献摘要

参考文献

被引文献

相似文献

交叉验证(CV)是估计分类器预测误差的一种有效方法。最近的一些文章提出了通过选择最小化CV误差估计的分类器参数值来优化分类器的方法。我们已经评估了使用优化分类器的CV误差估计作为独立数据上预期的真实误差估计的有效性。我们使用CV来优化两种分类器的分类参数:收缩质心和支持向量机(SVM)。创建随机训练数据集,两个类之间的特征分布没有差异。使用这些“空”数据集,我们选择了最小化CV误差估计的分类器参数值。10-倍CV用于收缩质心,而留一法CV(LOOCV)用于SVM。创建独立的测试数据以估计真实误差。对于“null”和“non null”(类之间具有差异表达)数据,我们还测试了嵌套CV程序,其中内部CV循环用于执行参数调整,而外部CV用于计算误差估计。发现具有最佳参数的分类器的CV误差估计是分类器在独立数据上会产生的真实误差的实质上有偏估计。即使对于“空”数据集,两个类之间没有真实的差异,但在18.5%的模拟训练数据集上,具有最佳参数的收缩质心的CV误差估计小于30%。对于最优参数的SVM,在38%的“空”数据集上的估计错误率小于30%。在独立测试集上,优化的分类器的性能并不比机会好。嵌套CV程序大大降低了偏差,并给出了一个估计的错误,这是非常接近的独立测试集上获得的收缩质心和SVM分类器的“空”和“非空”的数据分布。我们表明,使用CV来计算一个分类器,它本身已经调整使用CV的错误估计给出了一个显着的偏差估计的真实错误。正确使用CV来估计使用明确定义的算法开发的分类器的真实误差需要在每个CV循环中重复算法的所有步骤,包括分类器参数调整。嵌套CV程序提供了真实误差的几乎无偏估计。
Cross-validation (CV) is an effective method for estimating the prediction error of a classifier. Some recent articles have proposed methods for optimizing classifiers by choosing classifier parameter values that minimize the CV error estimate. We have evaluated the validity of using the CV error estimate of the optimized classifier as an estimate of the true error expected on independent data. We used CV to optimize the classification parameters for two kinds of classifiers; Shrunken Centroids and Support Vector Machines (SVM). Random training datasets were created, with no difference in the distribution of the features between the two classes. Using these "null" datasets, we selected classifier parameter values that minimized the CV error estimate. 10-fold CV was used for Shrunken Centroids while Leave-One-Out-CV (LOOCV) was used for the SVM. Independent test data was created to estimate the true error. With "null" and "non null" (with differential expression between the classes) data, we also tested a nested CV procedure, where an inner CV loop is used to perform the tuning of the parameters while an outer CV is used to compute an estimate of the error. The CV error estimate for the classifier with the optimal parameters was found to be a substantially biased estimate of the true error that the classifier would incur on independent data. Even though there is no real difference between the two classes for the "null" datasets, the CV error estimate for the Shrunken Centroid with the optimal parameters was less than 30% on 18.5% of simulated training data-sets. For SVM with optimal parameters the estimated error rate was less than 30% on 38% of "null" data-sets. Performance of the optimized classifiers on the independent test set was no better than chance. The nested CV procedure reduces the bias considerably and gives an estimate of the error that is very close to that obtained on the independent testing set for both Shrunken Centroids and SVM classifiers for "null" and "non-null" data distributions. We show that using CV to compute an error estimate for a classifier that has itself been tuned using CV gives a significantly biased estimate of the true error. Proper use of CV for estimating true error of a classifier developed using a well defined algorithm requires that all steps of the algorithm, including classifier parameter tuning, be repeated in each CV loop. A nested CV procedure provides an almost unbiased estimate of the true error.
DOI: 10.1093/bioinformatics/bti499
发表时间: 2005-08-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Molinaro, AM;Simon, R;Pfeiffer, RM
通讯作者: Pfeiffer, RM
DOI: 10.1093/bioinformatics/bti294
发表时间: 2005-05-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Fu, WJJ;Carroll, RJ;Wang, SJ
通讯作者: Wang, SJ
DOI: 10.1016/s0140-6736(03)12775-4
发表时间: 2003-03-15
期刊: LANCET
影响因子: 168.9
作者:
Iizuka, N;Oka, M;Hamamoto, Y
通讯作者: Hamamoto, Y
DOI: 10.1073/pnas.102102699
发表时间: 2002-05-14
影响因子: 11.1
作者:
Ambroise, C;McLachlan, GJ
通讯作者: McLachlan, GJ
DOI: 10.1073/pnas.082099299
发表时间: 2002-05-14
影响因子: 11.1
作者:
Tibshirani, R;Hastie, T;Chu, G
通讯作者: Chu, G