Reliable estimation of prediction errors for QSAR models under model uncertainty using double cross-validation.

Reliable estimation of prediction errors for QSAR models under model uncertainty using double cross-validation.
复制标题

DOI:
10.1186/s13321-014-0047-1
复制
发表时间:
2014
影响因子:
8.6
通讯作者:
Baumann K
Baumann K
中科院分区:
化学2区
文献类型:
--
作者:
Baumann D;Baumann K

文献摘要

参考文献

被引文献

相似文献

一般来说,QSAR建模需要模型选择和验证,因为没有关于最佳QSAR模型的先验知识。预测误差(PE)经常被用来选择和评估所研究的模型。预测误差的可靠估计具有挑战性-特别是在模型不确定性下-并且需要独立的测试对象。这些测试对象不得参与模型构建或模型选择。双重交叉验证,有时也被称为嵌套交叉验证,提供了一个有吸引力的可能性,以产生测试数据和选择QSAR模型,因为它使用的数据非常有效。然而,在模型不确定性下,双交叉验证的可靠性在文献中存在争议。此外,系统的研究调查充分参数化的双重交叉验证仍然缺失。在这里,交叉验证设计的内部循环和测试集大小的影响,在外部循环系统地研究了回归模型结合变量选择。通过双重交叉验证分析模拟和真实的数据,以识别影响所得模型质量的重要因素。对于模拟数据,提供了偏差-方差分解。结合变量选择的QSAR/QSPR回归模型的预测误差在很大程度上取决于双交叉验证的参数化。双交叉验证内环的参数主要影响模型的偏差和方差,外环的参数主要影响预测误差估计的变异性。双重交叉验证可靠和无偏地估计回归模型在模型不确定性下的预测误差。与单一测试集相比,双重交叉验证提供了更真实的模型质量,应优于单一测试集。本文的在线版本(doi:10.1186/s13321-014-0047-1)包含补充材料,可供授权用户使用。
Generally, QSAR modelling requires both model selection and validation since there is no a priori knowledge about the optimal QSAR model. Prediction errors (PE) are frequently used to select and to assess the models under study. Reliable estimation of prediction errors is challenging – especially under model uncertainty – and requires independent test objects. These test objects must not be involved in model building nor in model selection. Double cross-validation, sometimes also termed nested cross-validation, offers an attractive possibility to generate test data and to select QSAR models since it uses the data very efficiently. Nevertheless, there is a controversy in the literature with respect to the reliability of double cross-validation under model uncertainty. Moreover, systematic studies investigating the adequate parameterization of double cross-validation are still missing. Here, the cross-validation design in the inner loop and the influence of the test set size in the outer loop is systematically studied for regression models in combination with variable selection. Simulated and real data are analysed with double cross-validation to identify important factors for the resulting model quality. For the simulated data, a bias-variance decomposition is provided. The prediction errors of QSAR/QSPR regression models in combination with variable selection depend to a large degree on the parameterization of double cross-validation. While the parameters for the inner loop of double cross-validation mainly influence bias and variance of the resulting models, the parameters for the outer loop mainly influence the variability of the resulting prediction error estimate. Double cross-validation reliably and unbiasedly estimates prediction errors under model uncertainty for regression models. As compared to a single test set, double cross-validation provided a more realistic picture of model quality and should be preferred over a single test set. The online version of this article (doi:10.1186/s13321-014-0047-1) contains supplementary material, which is available to authorized users.
DOI: 10.1016/j.chemolab.2006.04.021
发表时间: 2006-12-01
影响因子: 3.9
作者:
Anderssen, Endre;Dyrstad, Knut;Martens, Harald
通讯作者: Martens, Harald
DOI: 10.1016/s0165-9936(03)00607-1
发表时间: 2003-06-01
影响因子: 13.1
作者:
Baumann, K
通讯作者: Baumann, K
DOI: 10.1186/1471-2105-9-360
发表时间: 2008-09-02
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Eklund, Martin;Spjuth, Ola;Wikberg, Jarl E. S.
通讯作者: Wikberg, Jarl E. S.
DOI: 10.1186/1471-2105-6-50
发表时间: 2005-03-10
期刊: BMC bioinformatics
影响因子: 3
作者:
Freyhult E;Prusis P;Lapinsh M;Wikberg JE;Moulton V;Gustafsson MG
通讯作者: Gustafsson MG
DOI: 10.1093/jnci/djj330
发表时间: 2006-09-06
影响因子: 10.3
作者:
Asgharzadeh, Shahab;Pique-Regi, Roger;Seeger, Robert C.
通讯作者: Seeger, Robert C.