Assessing model fit by cross-validation

Assessing model fit by cross-validation
复制标题

DOI:
10.1021/ci025626i
复制
发表时间:
2003-03-01
期刊:
JOURNAL OF CHEMICAL INFORMATION AND COMPUTER SCIENCES
影响因子:
--
通讯作者:
Mills, D
Mills, D
中科院分区:
其他
文献类型:
--
作者:
Hawkins, DM;Basak, SC;Mills, D

文献摘要

被引文献

相似文献

当QSAR模型拟合时,重要的是要验证任何拟合的模型,以检查其预测是否可能延续到模型拟合工作中未使用的新数据。有两种标准的方法可以做到这一点——使用一个单独的保留测试样本和计算上更繁琐的留一交叉验证,其中使用整个可用化合物池来拟合模型并评估其有效性。我们通过理论论证和对大型QSAR数据集的实证研究表明,当可用的样本量很小时——几十个或几十个而不是几百个——保留其中的一部分进行测试是浪费的,使用交叉验证要好得多,但要确保这是正确完成的。
When QSAR models are fitted, it is important to validate any fitted model-to check that it is plausible that its predictions will carry over to fresh data not used in the model fitting exercise. There are two standard ways of doing this-using a separate hold-out test sample and the computationally much more burdensome leave-one-out cross-validation in which the entire pool of available compounds is used both to fit the model and to assess its validity. We show by theoretical argument and empiric study of a large QSAR data set that when the available sample size is small-in the dozens or scores rather than the hundreds, holding a portion of it back for testing is wasteful, and that it is much better to use cross-validation, but ensure that this is done properly.