Structural and Sequence Similarity Makes a Significant Impact on Machine-Learning-Based Scoring Functions for Protein-Ligand Interactions

Structural and Sequence Similarity Makes a Significant Impact on Machine-Learning-Based Scoring Functions for Protein-Ligand Interactions
复制标题

结构和序列相似性对基于机器学习的蛋白质-配体相互作用评分函数产生重大影响

DOI:
10.1021/acs.jcim.7b00049
复制
发表时间:
2017-04-01
影响因子:
5.6
通讯作者:
Yang, Jianyi
Yang, Jianyi
中科院分区:
化学2区
文献类型:
--
作者:
Li, Yang;Yang, Jianyi

文献摘要

被引文献

相似文献

最近,基于机器学习的评分函数显著改善了蛋白质-配体结合亲和力的预测。例如,使用一组表示原子距离计数的简单描述符,RF分数在PDBbind 2007数据库的核心集上将Pearson相关系数提高到约0.8,这显著高于在相同基准上的任何常规评分函数的性能。已经进行了一些研究来讨论基于机器学习的方法的性能,但这种改进的原因仍然不清楚。在这项研究中,通过系统地控制PDBbind基准的训练和测试蛋白质之间的结构和序列相似性,我们证明了蛋白质结构和序列相似性对基于机器学习的方法产生了重大影响。在去除与通过结构比对和序列比对识别的测试蛋白高度相似的训练蛋白之后,在新的训练集上训练的基于机器学习的方法不再优于传统的评分函数。相反,像X-Score这样的传统函数的性能相对稳定,无论使用什么训练数据来拟合其能量项的权重。
The prediction of protein-ligand binding affinity has recently been improved remarkably by machine-learning-based scoring functions. For example, using a set Of simple descriptors representing the atomic distance counts, the RF-Score improves the Pearson correlation coefficient to about 0.8 on the core set of the PDBbind 2007 database, which is significantly higher than the performance of any conventional scoring function on the same benchmark. A few studies have been made to discuss the performance of machine-learning-based, methods, but the reason for this improvement remains unclear. In this study, by systemically controlling the structural and sequence similarity between the training and test proteins of the PDBbind benchmark, we demonstrate that protein structural and sequence Similarity makes a significant impact on machine-learning-based methods. After removal of training proteins that are highly similar to the test proteins identified by structure alignment and sequence alignment, machine-learning-based methods trained on the new training sets do not outperform the conventional scoring functions any more. On the contrary, the performance of conventional functions like X-Score is relatively stable no matter what training data are used to fit the weights of its energy terms.