Predicting functionally important residues from sequence conservation

Predicting functionally important residues from sequence conservation
复制标题

DOI:
10.1093/bioinformatics/btm270
复制
发表时间:
2007-08-01
期刊:
影响因子:
5.8
通讯作者:
Singh, Mona
Singh, Mona
中科院分区:
生物学3区
文献类型:
--
作者:
Capra, John A.;Singh, Mona

文献摘要

被引文献

相似文献

动机:蛋白质中的所有残基并不同等重要。有些是蛋白质的正确结构和功能所必需的,而另一些则很容易被取代。保守性分析是预测蛋白质序列中这些功能重要的残基的最广泛使用的方法之一。结果:我们介绍了一种基于詹森-香农散度估计序列保守性的信息论方法。我们还开发了一个一般的启发式,认为估计保守的顺序相邻的网站。在大规模测试中,我们证明了我们的组合方法在识别功能重要的残基方面优于以前的基于保护的措施;特别是,它明显优于常用的香农熵度量。我们发现,考虑顺序邻居的保守性可以提高所有测试方法的性能。我们的分析还表明,许多现有的方法,试图将氨基酸之间的关系不会导致更好地识别功能重要的网站。最后,我们发现,虽然保守是高度预测在确定催化位点和残基附近的结合配体,它是在确定蛋白质-蛋白质界面的残基有效得多。
Motivation: All residues in a protein are not equally important. Some are essential for the proper structure and function of the protein, whereas others can be readily replaced. Conservation analysis is one of the most widely used methods for predicting these functionally important residues in protein sequences.Results: We introduce an information-theoretic approach for estimating sequence conservation based on Jensen-Shannon divergence. We also develop a general heuristic that considers the estimated conservation of sequentially neighboring sites. In largescale testing, we demonstrate that our combined approach outperforms previous conservation-based measures in identifying functionally important residues; in particular, it is significantly better than the commonly used Shannon entropy measure. We find that considering conservation at sequential neighbors improves the performance of all methods tested. Our analysis also reveals that many existing methods that attempt to incorporate the relationships between amino acids do not lead to better identification of functionally important sites. Finally, we find that while conservation is highly predictive in identifying catalytic sites and residues near bound ligands, it is much less effective in identifying residues in protein-protein interfaces.