Understanding and identifying amino acid repeats

Understanding and identifying amino acid repeats
复制标题

DOI:
10.1093/bib/bbt003
复制
发表时间:
2014-07-01
影响因子:
9.5
通讯作者:
Nijveen, Harm
Nijveen, Harm
中科院分区:
生物学2区
文献类型:
--
作者:
Luo, Hong;Nijveen, Harm

文献摘要

被引文献

相似文献

氨基酸重复序列(阿尔斯)在蛋白质序列中含量丰富。它们在蛋白质功能和进化中具有特殊作用。DNA滑动产生的简单重复模式往往会在重复区域引入长度变异和点突变。由于其长度可变,正常功能的丧失和异常功能的获得是导致疾病的潜在风险。具有复杂模式的重复序列主要是指功能域重复序列,如著名的富含亮氨酸重复序列和WD重复序列,它们经常参与蛋白质-蛋白质相互作用。它们主要来源于内部基因复制事件,并由“看门人”残基稳定,这在防止结构域间聚集中起着至关重要的作用。阿尔斯广泛分布于不同的蛋白质组中,跨越各种分类学范围,并且在真核生物蛋白质中尤其丰富。然而,它们的具体演变和功能情景仍然知之甚少。识别蛋白质序列中的阿尔斯是深入研究其生物学功能和进化机制的第一步。原则上,这是一个NP难问题,因为大多数重复片段是由一系列复杂的进化事件形成的,并成为潜在的周期性模式。不可能定义用于检测和验证各种重复模式的统一标准。相反,已经开发了基于不同策略的不同算法来科普不同的重复模式。在这篇综述中,我们试图描述目前可用的氨基酸重复序列检测算法,并在深入分析蛋白质重复序列的生物学意义的基础上比较它们的策略。
Amino acid repeats (AARs) are abundant in protein sequences. They have particular roles in protein function and evolution. Simple repeat patterns generated by DNA slippage tend to introduce length variations and point mutations in repeat regions. Loss of normal and gain of abnormal function owing to their variable length are potential risks leading to diseases. Repeats with complex patterns mostly refer to the functional domain repeats, such as the well-known leucine-rich repeat and WD repeat, which are frequently involved in protein-protein interaction. They are mainly derived from internal gene duplication events and stabilized by 'gate-keeper' residues, which play crucial roles in preventing inter-domain aggregation. AARs are widely distributed in different proteomes across a variety of taxonomic ranges, and especially abundant in eukaryotic proteins. However, their specific evolutionary and functional scenarios are still poorly understood. Identifying AARs in protein sequences is the first step for the further investigation of their biological function and evolutionary mechanism. In principle, this is an NP-hard problem, as most of the repeat fragments are shaped by a series of sophisticated evolutionary events and become latent periodical patterns. It is not possible to define a uniform criterion for detecting and verifying various repeat patterns. Instead, different algorithms based on different strategies have been developed to cope with different repeat patterns. In this review, we attempt to describe the amino acid repeat-detection algorithms currently available and compare their strategies based on an in-depth analysis of the biological significance of protein repeats.