ATPbind: Accurate Protein-ATP Binding Site Prediction by Combining Sequence-Profiling and Structure-Based Comparisons.

ATPbind: Accurate Protein-ATP Binding Site Prediction by Combining Sequence-Profiling and Structure-Based Comparisons.
复制标题

ATPbind:通过结合序列分析和基于结构的比较来准确预测蛋白质-ATP 结合位点

DOI:
10.1021/acs.jcim.7b00397
复制
发表时间:
2018-02-26
影响因子:
5.6
通讯作者:
Yu DJ
Yu DJ
中科院分区:
化学2区
文献类型:
--
作者:
Hu J;Li Y;Zhang Y;Yu DJ

文献摘要

参考文献

被引文献

相似文献

蛋白质-ATP相互作用普遍存在于各种生物过程中。从蛋白质信息中准确定位ATP结合位点是蛋白质功能注释和药物发现的一项重要而具有挑战性的任务。然而,没有方法可以最佳地识别不同蛋白质的ATP结合位点。在这项研究中,我们报告了一个新的复合预测,ATP结合位点,通过整合两个基于模板的预测(即,S-SITE和TM-SITE)和蛋白质的三个区别性序列驱动特征:位置特异性评分矩阵、预测的二级结构和预测的溶剂可及性。在ATPbind中,我们组装多个支持向量机(SVMs)的基础上的随机欠采样技术,以科普ATP结合位点和非ATP结合位点的数量之间的严重失衡现象。我们还构建了一个新的黄金标准基准数据集,由PDB数据库中的429个ATP结合蛋白组成,以评估和比较拟议的ATP结合与其他现有的预测。从查询序列和预测的I-TASSER模型开始,ATPBind可以达到72%的平均准确度,覆盖所有ATP结合位点的62%,同时达到显著高于其他最先进预测器的马修斯相关系数值。
Protein–ATP interactions are ubiquitous in a wide variety of biological processes. Correctly locating ATP binding sites from protein information is an important but challenging task for protein function annotation and drug discovery. However, there is no method that can optimally identify ATP binding sites for different proteins. In this study, we report a new composite predictor, ATPbind, for ATP binding sites by integrating the outputs of two template-based predictors (i.e., S-SITE and TM-SITE) and three discriminative sequence-driven features of proteins: position specific scoring matrix, predicted secondary structure, and predicted solvent accessibility. In ATPbind, we assembled multiple support vector machines (SVMs) based on a random undersampling technique to cope with the serious imbalance phenomenon between the numbers of ATP binding sites and of non-ATP binding sites. We also constructed a new gold-standard benchmark data set consisting of 429 ATP binding proteins from the PDB database to evaluate and compare the proposed ATPbind with other existing predictors. Starting from a query sequence and predicted I-TASSER models, ATPbind can achieve an average accuracy of 72%, covering 62% of all ATP binding sites while achieving a Matthews correlation coefficient value that is significantly higher than that of other state-of-the-art predictors.
一种应用于蛋白质-核苷酸结合残基预测的新型监督过采样算法
DOI: 10.1371/journal.pone.0107676
发表时间: 2014
期刊: PloS one
影响因子: 3.7
作者:
Hu J;He X;Yu DJ;Yang XB;Yang JY;Shen HB
通讯作者: Shen HB
DOI: 10.1002/prot.24074
发表时间: 2012-07-01
影响因子: 2.9
作者:
Joo, Keehyoung;Lee, Sung Jong;Lee, Jooyoung
通讯作者: Lee, Jooyoung
基于KNN的动态查询驱动的类不平衡学习样本重缩放策略
DOI: 10.1016/j.neucom.2016.01.043
发表时间: 2016-05-26
期刊: NEUROCOMPUTING
影响因子: 6
作者:
Hu, Jun;Li, Yang;Yu, Dong-Jun
通讯作者: Yu, Dong-Jun
DOI: 10.1093/bioinformatics/bti315
发表时间: 2005-05-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Laurie, ATR;Jackson, RM
通讯作者: Jackson, RM
DOI: 10.1186/1471-2105-10-168
发表时间: 2009-06-02
期刊: BMC bioinformatics
影响因子: 3
作者:
Le Guilloux V;Schmidtke P;Tuffery P
通讯作者: Tuffery P