ATPsite: sequence-based prediction of ATP-binding residues.

ATPsite: sequence-based prediction of ATP-binding residues.
复制标题

DOI:
10.1186/1477-5956-9-s1-s4
复制
发表时间:
2011-10-14
期刊:
影响因子:
2
通讯作者:
Kurgan L
Kurgan L
中科院分区:
生物学4区
文献类型:
--
作者:
Chen K;Mizianty MJ;Kurgan L

文献摘要

被引文献

相似文献

ATP是一种普遍存在的核苷酸,为细胞活动提供能量,催化化学反应,并参与细胞信号传导。ATP-蛋白质相互作用的知识有助于蛋白质功能的注释,并在药物设计中找到应用。序列与结构注释的差距促使开发基于高通量序列的ATP结合残基预测因子。此外,我们的实证测试表明,唯一现有的预测,ATPint,其特征在于相对较低的预测质量。我们提出了一种新的,高通量的机器学习为基础的预测,ATP位点,它确定ATP结合残基的蛋白质序列。我们的预测器利用支持向量机分类器和一套全面的输入功能,这些功能是基于序列,进化概况和序列预测的结构描述符,包括二级结构,溶剂可及性和二面角。与现有的方法相比,ATPsite实现了显著更高的马修斯相关系数(MCC)和ROC曲线下面积(AUC)值,包括ATPint,基于保护率的rate 4site和基于保护率的BLAST预测。我们还评估了个别输入类型的有效性。PSSM配置文件,保护分数,和某些功能的基础上的氨基酸组被证明是更有效地预测ATP结合残基比其余的功能组。统计测试表明,ATPsite显着优于现有的解决方案。的共识的ATP位点与基于序列比对的预测显示,给进一步的改善。
ATP is a ubiquitous nucleotide that provides energy for cellular activities, catalyzes chemical reactions, and is involved in cellular signalling. The knowledge of the ATP-protein interactions helps with annotation of protein functions and finds applications in drug design. The sequence to structure annotation gap motivates development of high-throughput sequence-based predictors of the ATP-binding residues. Moreover, our empirical tests show that the only existing predictor, ATPint, is characterized by relatively low predictive quality. We propose a novel, high-throughput machine learning-based predictor, ATPsite, which identifies ATP-binding residues from protein sequences. Our predictor utilizes Support Vector Machine classifier and a comprehensive set of input features that are based on the sequence, evolutionary profiles, and the sequence-predicted structural descriptors including secondary structure, solvent accessibility, and dihedral angles. The ATPsite achieves significantly higher Mathews Correlation Coefficient (MCC) and Area Under the ROC Curve (AUC) values when compared with the existing methods including the ATPint, conservation-based rate4site, and alignment-based BLAST predictors. We also assessed the effectiveness of individual input types. The PSSM profile, the conservation scores, and certain features based on amino acid groups are shown to be more effective in predicting the ATP-binding residues than the remaining feature groups. Statistical tests show that ATPsite significantly outperforms existing solutions. The consensus of the ATPsite with the sequence-alignment based predictor is shown to give further improvements.