Learning from the ligand: using ligand-based features to improve binding affinity prediction

Learning from the ligand: using ligand-based features to improve binding affinity prediction
复制标题

DOI:
10.1093/bioinformatics/btz665
复制
发表时间:
2020-02-01
期刊:
影响因子:
5.8
通讯作者:
Morris, Garrett M.
Morris, Garrett M.
中科院分区:
生物学3区
文献类型:
--
作者:
Boyles, Fergus;Deane, Charlotte M.;Morris, Garrett M.

文献摘要

被引文献

相似文献

动机:发现蛋白质结合亲和力预测的机器学习评分功能始终超过经典评分功能。通用亲和力预测的基于结构的评分功能通常使用描述从蛋白质配体复合物获得的相互作用的功能,并且有关配体本身的化学或拓扑特性的有限信息。回报:我们证明机器学习评分功能的性能始终是通过包含不同的基于配体的特征来改善。例如,在PDBBIND 2007、2013和2016 Core集中,将RF-SCORE V3与RDKIT分子描述符相结合的随机森林(RF),最高为0.836、0.780和0.821,与0.836、0.780和0.821相结合,与0.836、0.780和0.821相结合。单独使用RF得分V3的功能时,0.746和0.814。不包括与训练集中测试集类似的蛋白质和/或配体对评分功能的性能有重大影响,但不会消除基于配体的特征的预测能力。此外,仅使用基于配体的特征的RF在类似于经典评分功能的水平上是可预测的,并且似乎正在预测配体对其蛋白质靶标的平均结合亲和力。
Motivation: Machine learning scoring functions for protein-ligand binding affinity prediction have been found to consistently outperform classical scoring functions. Structure-based scoring functions for universal affinity prediction typically use features describing interactions derived from the protein-ligand complex, with limited information about the chemical or topological properties of the ligand itself.Results: We demonstrate that the performance of machine learning scoring functions are consistently improved by the inclusion of diverse ligand-based features. For example, a Random Forest (RF) combining the features of RF-Score v3 with RDKit molecular descriptors achieved Pearson correlation coefficients of up to 0.836, 0.780 and 0.821 on the PDBbind 2007, 2013 and 2016 core sets, respectively, compared to 0.790, 0.746 and 0.814 when using the features of RF-Score v3 alone. Excluding proteins and/or ligands that are similar to those in the test sets from the training set has a significant effect on scoring function performance, but does not remove the predictive power of ligand-based features. Furthermore a RF using only ligand-based features is predictive at a level similar to classical scoring functions and it appears to be predicting the mean binding affinity of a ligand for its protein targets.