A machine learning-based method to improve docking scoring functions and its application to drug repurposing.

A machine learning-based method to improve docking scoring functions and its application to drug repurposing.
复制标题

DOI:
10.1021/ci100369f
复制
发表时间:
2011-02-28
影响因子:
5.6
通讯作者:
Bourne PE
Bourne PE
中科院分区:
化学2区
文献类型:
--
作者:
Kinnings SL;Liu N;Tonge PJ;Jackson RM;Xie L;Bourne PE

文献摘要

参考文献

被引文献

相似文献

众所周知,对接评分函数是结合亲和力的弱预测因子。它们通常为有助于总体能量得分的单个能量项分配一组公共权重,然而,这些权重应该是基因家族依赖的。此外,他们错误地认为,个别相互作用有助于对总的结合亲和力以相加的方式。实际上,非共价相互作用通常以非线性方式相互依赖。在本文中,我们展示了如何使用支持向量机(SVM),训练关联组的单个能量项检索从分子对接与已知的结合亲和力的每个化合物从高通量筛选实验,可以用来提高已知的结合亲和力和那些预测的对接程序eHiTS之间的相关性。我们构建了两个预测模型;一个是使用BindingDB中的IC 50值训练的回归模型,另一个是使用有用诱饵目录(DUD)中的活性和诱饵化合物训练的分类模型。此外,为了解决高通量筛选数据集中阴性数据过度代表的问题,我们设计了一个多平面SVM训练程序的分类模型。与原始eHiTS评分函数相比,两种SVM的性能都有所提高,这突出了在从其各个分量中获得整体能量分数时使用非线性方法的潜力。我们应用上述方法来训练用于结核分枝杆菌(M.tb)InhA的直接抑制剂的新评分函数。通过将配体结合位点比较与新的评分函数相结合,我们提出磷酸二酯酶抑制剂可以潜在地重新用于靶向结核分枝杆菌InhA。我们的方法可以应用于其他基因家族的目标结构和活性数据是可用的,在这里提出的工作中所示。
Docking scoring functions are notoriously weak predictors of binding affinity. They typically assign a common set of weights to the individual energy terms that contribute to the overall energy score, however, these weights should be gene family-dependent. In addition, they incorrectly assume that individual interactions contribute towards the total binding affinity in an additive manner. In reality, noncovalent interactions often depend on one another in a nonlinear manner. In this paper we show how the use of support vector machines (SVMs), trained by associating sets of individual energy terms retrieved from molecular docking with the known binding affinity of each compound from high-throughput screening experiments, can be used to improve the correlation between known binding affinities and those predicted by the docking program eHiTS. We construct two prediction models; a regression model trained using IC50 values from BindingDB, and a classification model trained using active and decoy compounds from the Directory of Useful Decoys (DUD). Moreover, to address the issue of overrepresentation of negative data in high-throughput screening data sets, we have designed a multiple-planar SVM training procedure for the classification model. The increased performance that both SVMs give when compared with the original eHiTS scoring function highlights the potential for using nonlinear methods when deriving overall energy scores from their individual components. We apply the above methodology to train a new scoring function for direct inhibitors of M.tuberculosis (M.tb) InhA. By combining ligand binding site comparison with the new scoring function, we propose that phosphodiesterase inhibitors can potentially be repurposed to target M.tb InhA. Our methodology may be applied to other gene families for which target structures and activity data are available, as demonstrated in the work presented here.
DOI: 10.1007/s10822-008-9189-4
发表时间: 2008-03-01
影响因子: 3.5
作者:
Irwin, John J.
通讯作者: Irwin, John J.
DOI: 10.1371/journal.pcbi.1000423
发表时间: 2009-07
影响因子: 4.3
作者:
Kinnings SL;Liu N;Buchmeier N;Tonge PJ;Xie L;Bourne PE
通讯作者: Bourne PE
DOI: 10.1016/j.jmb.2010.02.007
发表时间: 2010-04-09
影响因子: 5.6
作者:
Baum, Bernhard;Muley, Laveena;Klebe, Gerhard
通讯作者: Klebe, Gerhard
DOI: 10.1021/jp0217839
发表时间: 2003-09-04
影响因子: 3.3
作者:
Boresch, S;Tettinger, F;Karplus, M
通讯作者: Karplus, M
DOI: 10.1021/ci100244v
发表时间: 2010-10-25
影响因子: 5.6
作者:
Durrant, Jacob D.;McCammon, J. Andrew
通讯作者: McCammon, J. Andrew