Evolutionary approach to predicting the binding site residues of a protein from its primary sequence

Evolutionary approach to predicting the binding site residues of a protein from its primary sequence
复制标题

DOI:
10.1073/pnas.1102210108
复制
发表时间:
2011-03-29
影响因子:
11.1
通讯作者:
Li, Wen-Hsiung
Li, Wen-Hsiung
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Tseng, Yan Yuan;Li, Wen-Hsiung

文献摘要

被引文献

相似文献

蛋白质结合位点残基,特别是催化残基,在蛋白质功能中起着核心作用。由于非冗余蛋白质数据库中超过99%的相似蛋白质序列没有结构信息,因此需要开发从蛋白质的一级序列预测其结合位点残基的方法。这项任务是非常具有挑战性的,因为结合位点残基仅构成蛋白质的一小部分。然而,蛋白质的结合位点残基聚集在其功能口袋中,并且它们的空间模式在进化中趋于保守。为了利用这些进化和结构原则,我们构建了一个类似于50,000个模板的数据库(称为包含口袋的片段数据库),每个模板不仅包括包含功能口袋的序列片段,还包括口袋的结构属性。为了使用这个数据库,我们设计了一个模板匹配技术,称为残基匹配分析,并建立了一个标准,选择查询序列的模板。最后,我们开发了一个概率模型,用于分配空间分数的模板和查询序列之间的匹配的残基在局部比对使用一组选定的评分矩阵,并计算每个匹配的残基在查询序列中的结合可能性。从可能性,可以预测查询序列中的结合位点残基。为我们的方法开发了一个自动计算管道。性能评估表明,我们的方法达到了70%的精度预测结合位点残基在60%的灵敏度。
Protein binding site residues, especially catalytic residues, play a central role in protein function. Because more than 99% of the similar to 12 million protein sequences in the nonredundant protein database have no structural information, it is desirable to develop methods to predict the binding site residues of a protein from its primary sequence. This task is highly challenging, because the binding site residues constitute only a small portion of a protein. However, the binding site residues of a protein are clustered in its functional pocket(s), and their spatial patterns tend to be conserved in evolution. To take advantage of these evolutionary and structural principles, we constructed a database of similar to 50,000 templates (called the pocket-containing segment database), each of which includes not only a sequence segment that contains a functional pocket but also the structural attributes of the pocket. To use this database, we designed a template-matching technique, termed residue-matching profiling, and established a criterion for selecting templates for a query sequence. Finally, we developed a probabilistic model for assigning spatial scores to matched residues between the template and query sequence in local alignments using a set of selected scoring matrices and for computing the binding likelihood of each matched residue in the query sequence. From the likelihoods, one can predict the binding site residues in the query sequence. An automated computational pipeline was developed for our method. A performance evaluation shows that our method achieves a 70% precision in predicting binding site residues at 60% sensitivity.