Critical assessment of QSAR models of environmental toxicity against Tetrahymena pyriformis:: Focusing on applicability domain and overfitting by variable selection

Critical assessment of QSAR models of environmental toxicity against Tetrahymena pyriformis:: Focusing on applicability domain and overfitting by variable selection
复制标题

DOI:
10.1021/ci800151m
复制
发表时间:
2008-09-01
影响因子:
5.6
通讯作者:
Varnek, Alexandre
Varnek, Alexandre
中科院分区:
化学2区
文献类型:
--
作者:
Tetko, Igor V.;Sushko, Iurii;Varnek, Alexandre

文献摘要

被引文献

相似文献

预测精度的估计是QSAR建模中的一个关键问题。“到模型的距离”可以被定义为在特定模型的背景下针对给定性质定义训练集分子与测试集化合物之间的相似性的度量。它可以用许多不同的方式来表达,例如,使用Tanimoto系数,杠杆作用,在模型空间的相关性等,在本文中,我们已经使用了混合高斯分布以及统计测试,以评估六种类型的距离模型的能力,以区分化合物的小和大的预测误差。建立了12个水毒性的构效关系模型。梨形肌是用不同的机器学习方法和各种类型的描述符获得的。从模型集合计算的预测毒性的基于模型的油标准偏差的距离提供了最好的结果。这个距离也成功地区分分子与低和大的预测误差的机制为基础的模型开发使用日志P和最大受体超离域性描述符。因此,距离模型度量也可以用来增强机制QSAR模型,通过估计它们的预测误差。此外,预测的准确性主要取决于训练集数据在化学和活性空间中的分布,而不是用于开发模型的QSAR方法。我们已经证明,模型的不正确验证可能会导致对其性能的错误估计,并建议如何规避这个问题。分别对来自EPA高产量(HPV)挑战计划和EINECS(欧洲化学物质信息系统)的3182和48774分子的毒性进行了预测,并估计了预测的准确性。开发的模型可在http://www.qspr.org网站上在线获得。
The estimation of the accuracy of predictions is a critical problem in QSAR modeling. The "distance to model" can be defined as a metric that defines the similarity between the training set molecules and the test set compound for the given property in the context of a specific model. It could be expressed in many different ways, e.g., using Tanimoto coefficient, leverage, correlation in space of models, etc. In this paper we have used mixtures of Gaussian distributions as well as statistical tests to evaluate six types of distances to models with respect to their ability to discriminate compounds with small and large prediction errors. The analysis was performed for twelve QSAR models of aqueous toxicity against T. pyriformis obtained with different machine-learning methods and various types of descriptors. The distances to model based oil standard deviation of predicted toxicity calculated from the ensemble of models afforded the best results. This distance also successfully discriminated molecules with low and large prediction errors for a mechanism-based model developed using log P and the Maximum Acceptor Superdelocalizability descriptors. Thus, the distance to model metric could also be used to augment mechanistic QSAR models by estimating their prediction errors. Moreover, the accuracy of prediction is mainly determined by the training set data distribution in the chemistry and activity spaces but not by QSAR approaches used to develop the models. We have shown that incorrect validation of a model may result in the wrong estimation of its performance and suggested how this problem could be circumvented. The toxicity of 3182 and 48774 molecules from the EPA High Production Volume (HPV) Challenge Program and EINECS (European chemical Substances Information System), respectively, was predicted, and the accuracy of prediction was estimated. The developed models are available online at http://www.qspr.org site.