Correction to "Machine learning-based method to improve docking scoring functions and its application to drug repurposing".

Correction to "Machine learning-based method to improve docking scoring functions and its application to drug repurposing".
复制标题

DOI:
10.1021/ci2001346
复制
发表时间:
2011-05-23
影响因子:
5.6
通讯作者:
Bourne PE
Bourne PE
中科院分区:
化学2区
文献类型:
--
作者:
Kinnings SL;Liu N;Tonge PJ;Jackson RM;Xie L;Bourne PE

文献摘要

相似文献

众所周知,对接评分函数是结合亲和力的弱预测因子。他们通常为对整体能量得分做出贡献的各个能量项分配一组通用的权重;然而,这些权重应该取决于基因家族。此外,他们错误地假设个体相互作用以累加的方式对总结合亲和力做出贡献。实际上,非共价相互作用通常以非线性方式相互依赖。在本文中,我们展示了如何使用支持向量机(SVM)来改善已知结合亲和力与对接程序 eHiTS 预测的结合亲和力之间的相关性,该支持向量机(SVM)通过将分子对接中检索到的单个能量项集与高通量筛选实验中每种化合物的已知结合亲和力相关联来进行训练。我们构建了两个预测模型:使用 BindingDB 中的 IC50 值训练的回归模型,以及使用有用诱饵目录 (DUD) 中的活性和诱饵化合物训练的分类模型。此外,为了解决高通量筛选数据集中负面数据过多的问题,我们为分类模型设计了多平面支持向量机训练程序。与原始 eHiTS 评分函数相比,这两种 SVM 所提供的性能提高,凸显了在从其各个组件导出总体能量得分时使用非线性方法的潜力。我们应用上述方法来训练结核分枝杆菌(M.tb)InhA直接抑制剂的新评分函数。通过将配体结合位点比较与新的评分函数相结合,我们提出磷酸二酯酶抑制剂有可能重新用于靶向 M.tbInhA。我们的方法可以应用于可获得目标结构和活性数据的其他基因家族,如本文中介绍的工作所示。
Docking scoring functions are notoriously weak predictors of binding affinity. They typically assign a common set of weights to the individual energy terms that contribute to the overall energy score; however, these weights should be gene family dependent. In addition, they incorrectly assume that individual interactions contribute toward the total binding affinity in an additive manner. In reality, noncovalent interactions often depend on one another in a nonlinear manner. In this paper, we show how the use of support vector machines (SVMs), trained by associating sets of individual energy terms retrieved from molecular docking with the known binding affinity of each compound from high-throughput screening experiments, can be used to improve the correlation between known binding affinities and those predicted by the docking program eHiTS. We construct two prediction models: a regression model trained using IC50values from BindingDB, and a classification model trained using active and decoy compounds from the Directory of Useful Decoys (DUD). Moreover, to address the issue of overrepresentation of negative data in high-throughput screening data sets, we have designed a multiple-planar SVM training procedure for the classification model. The increased performance that both SVMs give when compared with the original eHiTS scoring function highlights the potential for using nonlinear methods when deriving overall energy scores from their individual components. We apply the above methodology to train a new scoring function for direct inhibitors ofMycobacterium tuberculosis(M.tb) InhA. By combining ligand binding site comparison with the new scoring function, we propose that phosphodiesterase inhibitors can potentially be repurposed to targetM.tbInhA. Our methodology may be applied to other gene families for which target structures and activity data are available, as demonstrated in the work presented here.