On Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation

On Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation
复制标题

DOI:
10.5555/1756006.1859921
复制
发表时间:
2010-03
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
G. Cawley;N. L. C. Talbot
G. Cawley;N. L. C. Talbot
中科院分区:
其他
文献类型:
--
作者:
G. Cawley;N. L. C. Talbot

文献摘要

被引文献

相似文献

机器学习算法的模型选择策略通常涉及对适当的模型选择标准进行数值优化,通常基于泛化性能的估计器,例如k重交叉验证。这种估计器的误差可以分解为偏差和方差分量。虽然无偏见经常被认为是模型选择标准的一个有益的品质,但我们证明了低方差至少同样重要,因为不可忽略的方差会在模型选择和训练模型时引入过度拟合的可能性。虽然这一观察结果在事后可能相当明显,但由于过度匹配模型选择标准而导致的性能降级可能会令人惊讶地大,这一观察结果似乎在机器学习文献中迄今几乎没有受到关注。在本文中,我们证明了这种形式的过拟合的影响通常与学习算法之间的性能差异具有相当的量级,因此在经验评估中不能被忽略。此外,我们还表明,由于这种形式的过度拟合,一些常见的绩效评估实践容易受到某种形式的选择偏差的影响,因此是不可靠的。我们讨论了避免模型选择中的过度拟合和绩效评估中随后的选择偏差的方法,我们希望这些方法将被纳入到最佳实践中。虽然这项研究集中在基于交叉验证的模型选择上,但研究结果相当普遍,适用于任何涉及对有限数据样本评估的模型选择标准进行优化的模型选择实践,包括最大化贝叶斯证据和优化性能界限。
Model selection strategies for machine learning algorithms typically involve the numerical optimisation of an appropriate model selection criterion, often based on an estimator of generalisation performance, such as k-fold cross-validation. The error of such an estimator can be broken down into bias and variance components. While unbiasedness is often cited as a beneficial quality of a model selection criterion, we demonstrate that a low variance is at least as important, as a non-negligible variance introduces the potential for over-fitting in model selection as well as in training the model. While this observation is in hindsight perhaps rather obvious, the degradation in performance due to over-fitting the model selection criterion can be surprisingly large, an observation that appears to have received little attention in the machine learning literature to date. In this paper, we show that the effects of this form of over-fitting are often of comparable magnitude to differences in performance between learning algorithms, and thus cannot be ignored in empirical evaluation. Furthermore, we show that some common performance evaluation practices are susceptible to a form of selection bias as a result of this form of over-fitting and hence are unreliable. We discuss methods to avoid over-fitting in model selection and subsequent selection bias in performance evaluation, which we hope will be incorporated into best practice. While this study concentrates on cross-validation based model selection, the findings are quite general and apply to any model selection practice involving the optimisation of a model selection criterion evaluated over a finite sample of data, including maximisation of the Bayesian evidence and optimisation of performance bounds.