Reducing over-optimism in variable selection by cross-model validation

Reducing over-optimism in variable selection by cross-model validation
复制标题

DOI:
10.1016/j.chemolab.2006.04.021
复制
发表时间:
2006-12-01
影响因子:
3.9
通讯作者:
Martens, Harald
Martens, Harald
中科院分区:
计算机科学3区
文献类型:
--
作者:
Anderssen, Endre;Dyrstad, Knut;Martens, Harald

文献摘要

被引文献

相似文献

对数学模型的拟合进行广泛的优化,以适应相对较小的一组经验数据,可能会导致过于乐观的验证结果。如果对最终优化模型的评估基于相同的验证方法和相同的输入数据,作为扩展模型优化的基础,则累积的虚假相关性可能在最终模型验证中显示为真实的预测能力。这方面的一个例子是在基于跨模型验证方案的多元回归中使用广泛的变量选择。为了说明基于传统的单层验证的优化中的过度乐观问题,将只包含随机数的人工数据集提交给回归建模。通过逐步变量选择对模型进行优化。在X变量从500个逐步减少到29个后,通过留一法交叉验证(84%),最终模型对X中的y有很好的明显预测能力。最后,在一个大的QSAR数据集上测试了跨模型验证的性能。随机选择了几个校准组,并通过变量选择优化了回归模型。将这些模型的预测精度与交叉验证和跨模型验证结果进行了比较。在这些测试中,跨模型验证可以更好地衡量模型预测能力。(C)2006年,爱思唯尔出版。
Extensive optimisation of a mathematical model's fit to a relatively small set of empirical data, may lead to over-optimistic validation results. If the assessment of the final, optimised model is based on the same validation method and the same input data that were used as basis for the extensive model optimisation, accumulated spurious correlations may appear as real predictive ability in the final model validation. An example of this is the use of extensive variable selection in multiple regression, based on a cross-model validation scheme.To illustrate the over-optimism problem in optimisation based on conventional one-layered validation, an artificial data set, with only random numbers was submitted to regression modelling. The model was optimised by stepwise variable selection. A very good apparent predictive ability for y from X was found in the final model by leave-one-out cross-validation (84%), after the number of X-variables had been reduced stepwise from 500 to 29. Finally, the performance of the cross-model validation is tested on one large QSAR data set. Several calibration sets were chosen randomly and a regression model optimised by variable selection. The prediction accuracy of these models was compared to the cross-validation and cross-model validation results. In these tests cross-model validation gives the better measure of model predictive ability. (c) 2006 Published by Elsevier B.V.