In Silico Prediction of Aqueous Solubility: A Multimodel Protocol Based on Chemical Similarity

In Silico Prediction of Aqueous Solubility: A Multimodel Protocol Based on Chemical Similarity
复制标题

DOI:
10.1021/mp300234q
复制
发表时间:
2012-11-01
影响因子:
4.9
通讯作者:
Miteva, Maria A.
Miteva, Maria A.
中科院分区:
医学2区
文献类型:
--
作者:
Chevillard, Florent;Lagorce, David;Miteva, Maria A.

文献摘要

被引文献

相似文献

水溶性是在药物发现过程中评估和优化的最重要的ADMET性质之一。目前,准确预测溶解度仍然非常具有挑战性,并且非常需要对现有的计算机模型进行独立的基准测试,例如提出改进的解决方案。在这项研究中,我们开发了一种新的协议,通过结合商业或免费软件包中现有的几个模型来改进溶解度预测。我们首先在几个数据集上对十个用于水溶性预测的计算机模型进行了评估,以评估方法的可靠性,并且我们提出了一个新的150个分子的多样化数据集作为相关测试集SolDiv150。我们开发了一个随机森林协议,以评估不同的指纹的水溶性预测基于分子结构相似性的性能。我们的协议,被称为“多模型协议”,允许选择最准确的模型之间所采用的模型或软件包的化合物的利益,实现r(2)的0.84时,应用到SolDiv150。我们还发现,这里评估的所有模型在类药物分子上的表现都比在真实的药物上的表现更好,因此在这个方向上需要进一步的改进。总的来说,我们的方法扩大了适用范围,如使用我们的协议相比,使用单独的模型获得的溶解度预测的更准确的结果所示。
Aqueous solubility is one of the most important ADMET properties to assess and to optimize during the drug discovery process. At present, accurate prediction of solubility remains very challenging and there is an important need of independent benchmarking of the existing in silico models such as to suggest solutions for their improvement. In this study, we developed a new protocol for improved solubility prediction by combining several existing models available in commercial or free software packages. We first performed an evaluation of ten in silico models for aqueous solubility prediction on several data sets in order to assess the reliability of the methods, and we proposed a new diverse data set of 150 molecules as relevant test set, SolDiv150. We developed a random forest protocol to evaluate the performance of different fingerprints for aqueous solubility prediction based on molecular structure similarity. Our protocol, called a "multimodel protocol", allows selecting the most accurate model for a compound of interest among the employed models or software packages, achieving r(2) of 0.84 when applied to SolDiv150. We also found that all models assessed here performed better on druglike molecules than on real drugs, thus additional improvement is needed in this direction. Overall, our approach enlarges the applicability domain as demonstrated by the more accurate results for solubility prediction obtained using our protocol in comparison to using individual models.