Protein-Ligand Docking Surrogate Models: A SARS-CoV-2 Benchmark for Deep Learning Accelerated Virtual Screening

Protein-Ligand Docking Surrogate Models: A SARS-CoV-2 Benchmark for Deep Learning Accelerated Virtual Screening
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Austin R. Clyde;T. Brettin;A. Partin;H. Yoo;Y. Babuji;B. Blaiszik;André Merzky;M. Turilli;S. Jha;A. Ramanathan;Rick L. Stevens
Austin R. Clyde;T. Brettin;A. Partin;H. Yoo;Y. Babuji;B. Blaiszik;André Merzky;M. Turilli;S. Jha;A. Ramanathan;Rick L. Stevens
中科院分区:
其他
文献类型:
--
作者:
Austin R. Clyde;T. Brettin;A. Partin;H. Yoo;Y. Babuji;B. Blaiszik;André Merzky;M. Turilli;S. Jha;A. Ramanathan;Rick L. Stevens

文献摘要

被引文献

相似文献

我们提出了一个基准来研究蛋白质配体对接的替代模型准确性。我们共享一个数据集,其中包含 2 亿个 3D 复杂结构和 2D 结构分数,涵盖 SARS-CoV-2 蛋白质组中超过 15 个受体或结合位点的 1300 万个“库存”分子。我们的工作表明,在相同的超级计算机节点类型上,代理对接模型的吞吐量比标准对接协议高出六个数量级。我们通过在一天之内针对 10 亿个分子运行每个目标(每 GPU 秒 50k 预测)来展示高速代理模型的强大功能。我们展示了利用代理 ML 模型作为预过滤器进行对接的工作流程。我们的工作流程筛选化合物库的速度比标准技术快十倍,检测潜在最佳得分 0.1% 的化合物的错误率低于 0.01%。我们对加速的分析解释了,为了在对接范式下筛选更多分子,另一个数量级的加速必须来自模型精度而不是计算速度(如果增加,将不再改变我们筛选分子的吞吐量)。我们相信,这是社区开始专注于提高替代模型准确性的有力证据,以提高筛选大规模化合物库的能力,速度比当前技术快 100 倍甚至 1000 倍。
We propose a benchmark to study surrogate model accuracy for protein-ligand docking. We share a dataset consisting of 200 million 3D complex structures and 2D structure scores across a consistent set of 13 million "in-stock" molecules over 15 receptors, or binding sites, across the SARS-CoV-2 proteome. Our work shows surrogate docking models have six orders of magnitude more throughput than standard docking protocols on the same supercomputer node types. We demonstrate the power of high-speed surrogate models by running each target against 1 billion molecules in under a day (50k predictions per GPU seconds). We showcase a workflow for docking utilizing surrogate ML models as a pre-filter. Our workflow is ten times faster at screening a library of compounds than the standard technique, with an error rate less than 0.01\% of detecting the underlying best scoring 0.1\% of compounds. Our analysis of the speedup explains that to screen more molecules under a docking paradigm, another order of magnitude speedup must come from model accuracy rather than computing speed (which, if increased, will not anymore alter our throughput to screen molecules). We believe this is strong evidence for the community to begin focusing on improving the accuracy of surrogate models to improve the ability to screen massive compound libraries 100x or even 1000x faster than current techniques.