Does a more precise chemical description of protein-ligand complexes lead to more accurate prediction of binding affinity?

Does a more precise chemical description of protein-ligand complexes lead to more accurate prediction of binding affinity?
复制标题

DOI:
10.1021/ci500091r
复制
发表时间:
2014-03-24
影响因子:
5.6
通讯作者:
Blundell TL
Blundell TL
中科院分区:
化学2区
文献类型:
--
作者:
Ballester PJ;Schreyer A;Blundell TL

文献摘要

参考文献

被引文献

相似文献

预测大量不同分子对一系列大分子靶标的结合亲和力是一项极具挑战性的任务。尝试这种计算预测的评分函数对于利用和分析对接的输出是必不可少的,对接又是基于结构的药物设计等问题中的重要工具。经典的评分函数假设一个预定的理论启发的函数形式的变量之间的关系,描述了实验确定或建模的蛋白质-配体复合物的结构和它的结合亲和力。这种方法的固有问题是难以明确地模拟分子间相互作用对结合亲和力的各种贡献。基于机器学习回归模型的新评分函数能够有效地利用大量实验数据并避免对预定函数形式的需求,已被证明优于广泛的最新评分函数在广泛使用的基准测试中。在这里,我们调查的复杂的化学描述的预测能力的得分函数使用系统的电池的数值实验的影响。后者产生了迄今为止最准确的基准评分函数。引人注目的是,我们还发现,更精确的蛋白质-配体复合物的化学描述通常不会导致更准确的结合亲和力预测。我们讨论了四个因素,可能有助于这一结果:建模假设,相互依赖的代表性和回归,数据限制到绑定状态,和构象异构性的数据。
Predicting the binding affinities of large sets of diverse molecules against a range of macromolecular targets is an extremely challenging task. The scoring functions that attempt such computational prediction are essential for exploiting and analyzing the outputs of docking, which is in turn an important tool in problems such as structure-based drug design. Classical scoring functions assume a predetermined theory-inspired functional form for the relationship between the variables that describe an experimentally determined or modeled structure of a protein–ligand complex and its binding affinity. The inherent problem of this approach is in the difficulty of explicitly modeling the various contributions of intermolecular interactions to binding affinity. New scoring functions based on machine-learning regression models, which are able to exploit effectively much larger amounts of experimental data and circumvent the need for a predetermined functional form, have already been shown to outperform a broad range of state-of-the-art scoring functions in a widely used benchmark. Here, we investigate the impact of the chemical description of the complex on the predictive power of the resulting scoring function using a systematic battery of numerical experiments. The latter resulted in the most accurate scoring function to date on the benchmark. Strikingly, we also found that a more precise chemical description of the protein–ligand complex does not generally lead to a more accurate prediction of binding affinity. We discuss four factors that may contribute to this result: modeling assumptions, codependence of representation and regression, data restricted to the bound state, and conformational heterogeneity in data.
DOI: 10.1021/ci2003889
发表时间: 2011-11-28
影响因子: 5.6
作者:
Durrant JD;McCammon JA
通讯作者: McCammon JA
DOI: 10.1021/ci200377u
发表时间: 2011-12-27
影响因子: 5.6
作者:
Fan H;Schneidman-Duhovny D;Irwin JJ;Dong G;Shoichet BK;Sali A
通讯作者: Sali A
DOI: 10.1016/j.sbi.2008.11.009
发表时间: 2009-02
影响因子: 6.8
作者:
Guvench, Olgun;MacKerell, Alexander D., Jr.
通讯作者: MacKerell, Alexander D., Jr.
DOI: 10.1016/1074-5521(95)90050-0
发表时间: 1995-05-01
影响因子: --
作者:
GEHLHAAR, DK;VERKHIVKER, GM;FREER, ST
通讯作者: FREER, ST
DOI: 10.1016/j.jmb.2010.02.007
发表时间: 2010-04-09
影响因子: 5.6
作者:
Baum, Bernhard;Muley, Laveena;Klebe, Gerhard
通讯作者: Klebe, Gerhard