Task-Specific Scoring Functions for Predicting Ligand Binding Poses and Affinity and for Screening Enrichment

Task-Specific Scoring Functions for Predicting Ligand Binding Poses and Affinity and for Screening Enrichment
复制标题

DOI:
10.1021/acs.jcim.7b00309
复制
发表时间:
2018-01-01
影响因子:
5.6
通讯作者:
Mahapatra, Nihar R.
Mahapatra, Nihar R.
中科院分区:
化学2区
文献类型:
--
作者:
Ashtawy, Hossam M.;Mahapatra, Nihar R.

文献摘要

被引文献

相似文献

分子对接、评分和虚拟筛选在计算机辅助药物发现中发挥着越来越重要的作用。评分函数(SF)通常用于预测配体针对疾病途径中的关键蛋白质靶标的结合构象(对接任务)、结合亲和力(评分任务)和二元活性水平(筛选任务)。在当今可用的大多数分子对接软件包中,通用的基于结合亲和力(基于BA)的SF被用于所有三个任务以解决三个不同但相关的预测问题。在这三项任务中,这种SF的有限预测准确性一直是成本效益药物发现的主要障碍。因此,在这项工作中,我们开发了BT-Score,这是一种由提升决策树和数千个预测描述符组成的集成机器学习(ML)SF来估计BA。BT评分再现了样本外测试复合物的BA,相关性为0.825。即使在评分任务中具有如此高的准确性,我们也证明了BT分数和其他基于BA的SF的对接和筛选性能远非理想。这促使我们为对接和筛选问题构建两个特定于任务的ML SF。我们建议BT-Dock,a.增强树集成模型在大量的本地和计算机生成的配体构象上训练,并优化以明确预测结合姿势。该模型在不同的配体位姿预测方案中显示出比其基于BA的对应物平均提高25%。我们的基于筛选的SF,BT-Screen也获得了类似的改进,它直接将配体活性标记任务建模为分类问题。BT-Screen在数千种活性和非活性蛋白质-配体复合物上进行训练,以优化它,从其训练集中未发现的配体数据库中找到真实的活性物。除了这三个任务特定的SF之外,我们还提出了一种新的多任务深度神经网络(MT-Net),该网络在三个任务的数据上进行训练,以同时预测结合姿势,亲和力和活动水平。我们表明,MT-Net的性能是上级传统的SF和等同于或优于基于单任务神经网络的模型。
Molecular docking, scoring, and virtual screening play an increasingly important role in computer-aided drug discovery. Scoring functions (SFs) are typically, employed to predict the binding conformation (docking task), binding affinity (scoring task), and binary activity level (screening task) of ligands against a critical protein target in a disease's pathway. In most molecular docking software packages available today, a generic binding affinity-based (BA-based) SF is invoiced for all three tasks to solve three different, but related; prediction problems. The limited predictive accuracies of such SFs in these three tasks has been a major roadblock toward cost-effective drug discovery. Therefore, in this work, we develop BT-Score, an ensemble machine-learning (ML) SF of boosted decision trees and thousands Of predictive descriptors to estimate BA. BT-Score reproduced BA of out-of-sample test complexes with correlation of 0.825. Even with this high accuracy in the scoring task, we demonstrate that the docking and screening performance of BT-Score and other BA-based SFs is far from ideal. This has motivated us to build two task-specific ML SFs for the docking and screening problems. We propose BT-Dock, a. boosted-tree ensemble model trained on a large-number of native and computer-generated ligand conformations and optimized to predict binding poses explicitly. This model has shown an average improvement of 25% over its BA-based counterparts in different ligand pose prediction scenarios. Similar improvement has also been obtained by our screening-based SF, BT-Screen, which directly models the ligand activity labeling task as a classification problem. BT-Screen is trained-on thousands of active and inactive protein-ligand complexes to optimize it for finding real actives from databases of ligands not seen in its training set. In addition to the three task-specific SFs, we propose a novel multi-task deep neural network (MT-Net) that is trained on data from the three tasks to simultaneously predict binding poses, affinities, and activity levels. We show that the performance of MT-Net is superior to conventional SFs and on a par with or better than models based on single task neural networks.