Sensitivity Analysis of k-Fold Cross Validation in Prediction Error Estimation

Sensitivity Analysis of k-Fold Cross Validation in Prediction Error Estimation
复制标题

DOI:
10.1109/tpami.2009.187
复制
发表时间:
2010-03-01
影响因子:
23.6
通讯作者:
Antonio Lozano, Jose
Antonio Lozano, Jose
中科院分区:
计算机科学1区
文献类型:
--
作者:
Diego Rodriguez, Juan;Perez, Aritz;Antonio Lozano, Jose

文献摘要

被引文献

相似文献

在机器学习领域,分类器的性能通常以预测误差来衡量。在大多数现实世界的问题中,误差无法精确计算,必须估计。因此,选择适当的误差估计量是很重要的。本文分析了k折交叉验证分类误差估计量(k-cv)的统计性质,即偏差和方差。我们的主要贡献是一个新的理论分解的方差的k-cv考虑其来源的方差:敏感性的变化,在训练集和敏感性的变化的倍。本文还比较了不同k值下估计量的偏差和方差。实验研究已经在人工域中进行,因为它们允许精确计算隐含的量,并且我们可以严格指定实验条件。实验已经进行了两个分类器(朴素贝叶斯和最近邻),不同的折叠次数,样本大小,和训练集来自各种概率分布。最后,我们包括一些实用的建议,使用k折交叉验证。
In the machine learning field, the performance of a classifier is usually measured in terms of prediction error. In most real-world problems, the error cannot be exactly calculated and it must be estimated. Therefore, it is important to choose an appropriate estimator of the error. This paper analyzes the statistical properties, bias and variance, of the k-fold cross-validation classification error estimator (k-cv). Our main contribution is a novel theoretical decomposition of the variance of the k-cv considering its sources of variance: sensitivity to changes in the training set and sensitivity to changes in the folds. The paper also compares the bias and variance of the estimator for different values of k. The experimental study has been performed in artificial domains because they allow the exact computation of the implied quantities and we can rigorously specify the conditions of experimentation. The experimentation has been performed for two classifiers (naive Bayes and nearest neighbor), different numbers of folds, sample sizes, and training sets coming from assorted probability distributions. We conclude by including some practical recommendation on the use of k-fold cross validation.