Using Open Source Computational Tools for Predicting Human Metabolic Stability and Additional Absorption, Distribution, Metabolism, Excretion, and Toxicity Properties

Using Open Source Computational Tools for Predicting Human Metabolic Stability and Additional Absorption, Distribution, Metabolism, Excretion, and Toxicity Properties
复制标题

DOI:
10.1124/dmd.110.034918
复制
发表时间:
2010-11-01
影响因子:
3.9
通讯作者:
Ekins, Sean
Ekins, Sean
中科院分区:
医学2区
文献类型:
--
作者:
Gupta, Rishi R.;Gifford, Eric M.;Ekins, Sean

文献摘要

被引文献

相似文献

基于配体的计算模型可以更容易地在研究人员和组织之间共享,如果它们是用开源分子描述符生成的。例如,在一个实施例中,化学开发工具包(CDK)]和建模算法,因为这将否定对专有商业软件的要求。我们最初使用大约50,000个分子的训练集和大约25,000个分子的测试集以及人肝微粒体代谢稳定性数据评估了开源描述符和模型构建算法。C5.0决策树模型表明,CDK描述符与一组Smiles任意目标规范(SMARTS)键一起具有良好的统计学[kappa = 0.43,灵敏度= 0.57,特异性= 0.91,阳性预测值(PPV)= 0.64],等同于使用商用分子操作环境2D(MOE 2D)和相同的SMARTS密钥集构建的模型(kappa = 0.43,灵敏度= 0.58,特异性= 0.91,PPV = 0.63)。将数据集扩展到类似于193,000个分子,并使用Cubist与CDK和SMARTS键或MOE 2D和SMARTS键的组合生成连续模型,证实了这一观察结果。当连续预测值和实际值进行分组以获得分类评分时,我们观察到类似的kappa统计量(0.42)。将相同的描述符集和建模方法组合应用于具有相似模型检验统计量的被动渗透性和P-糖蛋白外排数据。总之,开放源码工具展示了与商业软件相当的预测结果,并节省了相应的成本。我们讨论了开源描述符的优点和缺点,以及它们作为组织竞争前共享数据的工具的机会,避免重复和协助药物发现。
Ligand-based computational models could be more readily shared between researchers and organizations if they were generated with open source molecular descriptors [e. g., chemistry development kit (CDK)] and modeling algorithms, because this would negate the requirement for proprietary commercial software. We initially evaluated open source descriptors and model building algorithms using a training set of approximately 50,000 molecules and a test set of approximately 25,000 molecules with human liver microsomal metabolic stability data. A C5.0 decision tree model demonstrated that CDK descriptors together with a set of Smiles Arbitrary Target Specification (SMARTS) keys had good statistics [kappa = 0.43, sensitivity = 0.57, specificity = 0.91, and positive predicted value (PPV) = 0.64], equivalent to those of models built with commercial Molecular Operating Environment 2D (MOE2D) and the same set of SMARTS keys (kappa = 0.43, sensitivity = 0.58, specificity = 0.91, and PPV = 0.63). Extending the dataset to similar to 193,000 molecules and generating a continuous model using Cubist with a combination of CDK and SMARTS keys or MOE2D and SMARTS keys confirmed this observation. When the continuous predictions and actual values were binned to get a categorical score we observed a similar kappa statistic (0.42). The same combination of descriptor set and modeling method was applied to passive permeability and P-glycoprotein efflux data with similar model testing statistics. In summary, open source tools demonstrated predictive results comparable to those of commercial software with attendant cost savings. We discuss the advantages and disadvantages of open source descriptors and the opportunity for their use as a tool for organizations to share data precompetitively, avoiding repetition and assisting drug discovery.