Three-Dimensional Convolutional Neural Networks and a Cross-Docked Data Set for Structure-Based Drug Design.

Three-Dimensional Convolutional Neural Networks and a Cross-Docked Data Set for Structure-Based Drug Design.
复制标题

DOI:
10.1021/acs.jcim.0c00411
复制
发表时间:
2020-09-28
影响因子:
5.6
通讯作者:
Koes DR
Koes DR
中科院分区:
化学2区
文献类型:
--
作者:
Francoeur PG;Masuda T;Sunseri J;Jia A;Iovanisci RB;Snyder I;Koes DR

文献摘要

参考文献

被引文献

相似文献

药物发现的主要挑战之一是预测蛋白质与配体的结合亲和力。最近,机器学习方法在这一任务上取得了实质性进展。然而,目前的模型评估方法在衡量对新目标的泛化方面过于乐观,并且没有足够大的标准数据集来比较模型之间的性能。我们提出了一个新的基于结构的机器学习数据集CrossDocked2020 Set,将2250万个配体的姿势对接到蛋白质数据库中的多个类似的结合口袋中,并在该数据集上对基于网格的卷积神经网络(CNN)模型进行了综合评估。我们还演示了训练数据和测试数据的划分如何影响使用PDBind数据集训练的模型的结果,如何通过添加更多质量较低的训练数据来提高性能,以及具有停靠姿势的训练如何对预测的复合体亲和力赋予姿势敏感性。我们最好的模型是一个由五个紧密连接的CNN组成的集成模型,在亲和力预测任务上的均方根误差为1.42%,皮尔逊回归系数为0.612,在绑定姿势分类时的AUC为0.956,在CrossDocked2020集合上的姿势选择准确率为68.4%。通过提供用于集群交叉验证的数据拆分和CrossDocked2020集合的原始数据,我们建立了第一个用于训练机器学习模型的标准化数据集,以识别非同源目标结构中的配体,同时也大大扩展了可用于训练的姿势数量。为了促进社区采用该数据集作为基准的蛋白质-配体结合亲和力预测,我们在https://github.com/gnina/models.提供了我们的模型、权重和CrosDocked2020集合
One of the main challenges in drug discovery is predicting protein-ligand binding affinity. Recently, machine learning approaches have made substantial progress on this task. However, current methods of model evaluation are overly optimistic in measuring generalization to new targets, and there does not exist a standard dataset of sufficient size to compare performance between models. We present a new dataset for structure-based machine learning, the CrossDocked2020 set, with 22.5 million poses of ligands docked into multiple similar binding pockets across the Protein Data Bank, and perform a comprehensive evaluation of grid-based convolutional neural network (CNN) models on this dataset. We also demonstrate how the partitioning of the training data and test data can impact the results of models trained with the PDBbind dataset, how performance improves by adding more lower-quality training data, and how training with docked poses imparts pose sensitivity to the predicted affinity of a complex. Our best performing model, an ensemble of five densely connected CNNs, achieves a root mean squared error of 1.42 and Pearson R of 0.612 on the affinity prediction task, an AUC of 0.956 at binding pose classification, and a 68.4% accuracy at pose selection on the CrossDocked2020 set. By providing data splits for clustered cross-validation and the raw data for the CrossDocked2020 set, we establish the first standardized dataset for training machine learning models to recognize ligands in non-cognate target structures while also greatly expanding the number of poses available for training. In order to facilitate community adoption of this dataset for benchmarking protein-ligand binding affinity prediction, we provide our models, weights, and the CrossDocked2020 set at https://github.com/gnina/models.
DOI: 10.1021/acs.jcim.7b00309
发表时间: 2018-01-01
影响因子: 5.6
作者:
Ashtawy, Hossam M.;Mahapatra, Nihar R.
通讯作者: Mahapatra, Nihar R.
DOI: 10.1021/ci2003889
发表时间: 2011-11-28
影响因子: 5.6
作者:
Durrant JD;McCammon JA
通讯作者: McCammon JA
DOI: 10.1021/ci400025f
发表时间: 2013-08-26
影响因子: 5.6
作者:
Damm-Ganamet KL;Smith RD;Dunbar JB Jr;Stuckey JA;Carlson HA
通讯作者: Carlson HA
DOI: 10.1371/journal.pone.0220113
发表时间: 2019-08-20
期刊: PLOS ONE
影响因子: 3.7
作者:
Chen, Lieyang;Cruz, Anthony;Kurtzman, Tom
通讯作者: Kurtzman, Tom
DOI: 10.1002/prot.21214
发表时间: 2007-02-01
影响因子: 2.9
作者:
Huang, Sheng-You;Zou, Xiaoqin
通讯作者: Zou, Xiaoqin