Cross-validation pitfalls when selecting and assessing regression and classification models.

Cross-validation pitfalls when selecting and assessing regression and classification models.
复制标题

DOI:
10.1186/1758-2946-6-10
复制
发表时间:
2014-03-29
影响因子:
8.6
通讯作者:
Thomas S
Thomas S
中科院分区:
化学2区
文献类型:
--
作者:
Krstajic D;Buturovic LJ;Leahy DE;Thomas S

文献摘要

参考文献

被引文献

相似文献

我们使用交叉验证来解决选择和评估分类和回归模型的问题。目前最先进的方法可以产生具有高方差的模型,使得它们不适合包括QSAR在内的许多实际应用。在本文中,我们描述和评估的最佳做法,提高可靠性,增加选定的模型的信心。所提出的方法的一个关键业务组成部分是云计算,它使以前不可行的方法的日常使用。我们详细描述了一种用于分类和回归中参数调整的重复网格搜索V折交叉验证算法,并定义了一种用于模型评估的重复嵌套交叉验证算法。至于变量的选择和参数调整,我们定义了两种算法(重复网格搜索交叉验证和双重交叉验证),并提供参数使用重复网格搜索在一般情况下。我们展示了我们的算法在七个QSAR数据集上的结果。在选择和评估分类和回归模型时,需要考虑预测性能的变化,这是在V折交叉验证中选择不同数据集分割的结果。我们证明了在选择最佳模型时重复交叉验证的重要性,以及在评估预测误差时重复嵌套交叉验证的重要性。本文的在线版本(doi:10.1186/1758-2946-6-10)包含补充材料,可供授权用户使用。
We address the problem of selecting and assessing classification and regression models using cross-validation. Current state-of-the-art methods can yield models with high variance, rendering them unsuitable for a number of practical applications including QSAR. In this paper we describe and evaluate best practices which improve reliability and increase confidence in selected models. A key operational component of the proposed methods is cloud computing which enables routine use of previously infeasible approaches. We describe in detail an algorithm for repeated grid-search V-fold cross-validation for parameter tuning in classification and regression, and we define a repeated nested cross-validation algorithm for model assessment. As regards variable selection and parameter tuning we define two algorithms (repeated grid-search cross-validation and double cross-validation), and provide arguments for using the repeated grid-search in the general case. We show results of our algorithms on seven QSAR datasets. The variation of the prediction performance, which is the result of choosing different splits of the dataset in V-fold cross-validation, needs to be taken into account when selecting and assessing classification and regression models. We demonstrate the importance of repeating cross-validation when selecting an optimal model, as well as the importance of repeating nested cross-validation when assessing a prediction error. The online version of this article (doi:10.1186/1758-2946-6-10) contains supplementary material, which is available to authorized users.
DOI: 10.1021/jm040835a
发表时间: 2005-01-13
影响因子: 7.3
作者:
Kazius, J;McGuire, R;Bursi, R
通讯作者: Bursi, R
DOI: 10.1214/09-ss054
发表时间: 2010-01-01
期刊: STATISTICS SURVEYS
影响因子: 3.3
作者:
Arlot, Sylvain;Celisse, Alain
通讯作者: Celisse, Alain
DOI: 10.1021/ci0500132
发表时间: 2005-05-01
影响因子: 5.6
作者:
Karthikeyan, M;Glen, RC;Bender, A
通讯作者: Bender, A
DOI: 10.1080/03610927608827333
发表时间: 1976-01-01
期刊: COMMUNICATIONS IN STATISTICS PART A-THEORY AND METHODS
影响因子: --
作者:
HOERL, AE;KENNARD, RW
通讯作者: KENNARD, RW
DOI: 10.1021/ci400113t
发表时间: 2013-06-01
影响因子: 5.6
作者:
Goracci, Laura;Ceccarelli, Martina;Cruciani, Gabriele
通讯作者: Cruciani, Gabriele