Deep Confident Steps to New Pockets: Strategies for Docking Generalization

Deep Confident Steps to New Pockets: Strategies for Docking Generalization
复制标题

DOI:
10.48550/arxiv.2402.18396
复制
发表时间:
2024-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Gabriele Corso;Arthur Deng;Benjamin Fry;Nicholas Polizzi;R. Barzilay;T. Jaakkola
Gabriele Corso;Arthur Deng;Benjamin Fry;Nicholas Polizzi;R. Barzilay;T. Jaakkola
中科院分区:
其他
文献类型:
--
作者:
Gabriele Corso;Arthur Deng;Benjamin Fry;Nicholas Polizzi;R. Barzilay;T. Jaakkola

文献摘要

相似文献

精确的盲对接有可能带来新的生物学突破,但为了实现这一承诺,对接方法必须在蛋白质组中得到很好的推广。然而,现有的基准未能严格评估普适性。因此,我们开发了一种新的基于蛋白质配体结合域的基准模型DockGen,并表明现有的基于机器学习的对接模型具有很弱的泛化能力。我们仔细分析了基于ML的对接的扩展规律,并表明,通过扩展数据和模型大小,以及集成合成数据策略,我们能够显著提高泛化能力,并在基准测试中设置新的最先进的性能。进一步,我们提出了置信度自举,这是一种新的训练范式,它完全依赖于扩散模型和置信度模型之间的相互作用,并利用了扩散模型的多分辨率生成过程。我们证明了置信度自举显著提高了基于ML的对接方法对接到未知蛋白质类的能力,更接近于准确和可推广的盲对接方法。
Accurate blind docking has the potential to lead to new biological breakthroughs, but for this promise to be realized, docking methods must generalize well across the proteome. Existing benchmarks, however, fail to rigorously assess generalizability. Therefore, we develop DockGen, a new benchmark based on the ligand-binding domains of proteins, and we show that existing machine learning-based docking models have very weak generalization abilities. We carefully analyze the scaling laws of ML-based docking and show that, by scaling data and model size, as well as integrating synthetic data strategies, we are able to significantly increase the generalization capacity and set new state-of-the-art performance across benchmarks. Further, we propose Confidence Bootstrapping, a new training paradigm that solely relies on the interaction between diffusion and confidence models and exploits the multi-resolution generation process of diffusion models. We demonstrate that Confidence Bootstrapping significantly improves the ability of ML-based docking methods to dock to unseen protein classes, edging closer to accurate and generalizable blind docking methods.