De novo identification of highly diverged protein repeats by probabilistic consistency

De novo identification of highly diverged protein repeats by probabilistic consistency
复制标题

DOI:
10.1093/bioinformatics/btn039
复制
发表时间:
2008-03-01
期刊:
影响因子:
5.8
通讯作者:
Soeding, J.
Soeding, J.
中科院分区:
生物学3区
文献类型:
--
作者:
Biegert, A.;Soeding, J.

文献摘要

被引文献

相似文献

动机:估计25%的真核蛋白质含有重复序列,这强调了重复对于进化新的蛋白质功能的重要性。内部重复序列通常与蛋白质的结构或功能单位相对应。因此,能够在序列水平上识别分散的重复片段或结构域的方法可以帮助预测结构域结构,推断关于功能和机制的假设,以及研究蛋白质从较小片段的进化。结果:我们提出了一种重新鉴定蛋白质序列重复序列的方法HHrepID。它能够检测许多尚未知道具有内部序列对称性的蛋白质的结构重复序列特征,例如外膜β -桶。HHrepID使用HMMHMM比较以同源物的多序列比对的形式挖掘进化信息。与之前的方法相比,新方法(1)生成重复序列的多重比对;(2)利用同源的传递性,通过一个新的合并程序与全概率处理对齐;(3)通过最大化期望精度的算法提高对准质量;(4)能够通过概率域边界检测方法识别复杂结构中不同类型的重复序列;(5)通过评估统计显著性的新方法提高灵敏度。
Motivation: An estimated 25% of all eukaryotic proteins contain repeats, which underlines the importance of duplication for evolving new protein functions. Internal repeats often correspond to structural or functional units in proteins. Methods capable of identifying diverged repeated segments or domains at the sequence level can therefore assist in predicting domain structures, inferring hypotheses about function and mechanism, and investigating the evolution of proteins from smaller fragments.Results: We present HHrepID, a method for the de novo identification of repeats in protein sequences. It is able to detect the sequence signature of structural repeats in many proteins that have not yet been known to possess internal sequence symmetry, such as outer membrane beta-barrels. HHrepID uses HMMHMM comparison to exploit evolutionary information in the form of multiple sequence alignments of homologs. In contrast to a previous method, the new method (1) generates a multiple alignment of repeats; (2) utilizes the transitive nature of homology through a novel merging procedure with fully probabilistic treatment of alignments; (3) improves alignment quality through an algorithm that maximizes the expected accuracy; (4) is able to identify different kinds of repeats within complex architectures by a probabilistic domain boundary detection method and (5) improves sensitivity through a new approach to assess statistical significance.