Mining protein loops using a structural alphabet and statistical exceptionality.

Mining protein loops using a structural alphabet and statistical exceptionality.
复制标题

DOI:
10.1186/1471-2105-11-75
复制
发表时间:
2010-02-04
期刊:
影响因子:
3
通讯作者:
Camproux AC
Camproux AC
中科院分区:
生物学4区
文献类型:
--
作者:
Regad L;Martin J;Nuel G;Camproux AC

文献摘要

参考文献

被引文献

相似文献

蛋白质环包含可用三维结构中 50% 的蛋白质残基。这些区域通常涉及蛋白质功能,例如结合位点、催化口袋……然而,用传统工具描述蛋白质环是一项艰巨的任务。规则的二级结构、螺旋和链已被广泛研究,而环由于其序列和结构变化很大,因此很难分析。由于数据稀疏,长循环很少被系统地研究。我们开发了一种简单而准确的方法,可以使用结构基序描述和分析短环和长环的结构,而不受环长度的限制。该方法基于结构字母表HMM-SA。 HMM-SA 允许将三维蛋白质结构简化为一维状态串,其中每个状态是一个四残基原型片段,称为结构字母。因此,通过像传统蛋白质序列分析中那样处理结构字母串,可以轻松完成对庞大数据集进行结构分组的艰巨任务。我们系统地提取了 93000 个蛋白质环库中的所有七残基片段,并根据结构字母序列对它们进行分组,称为结构词。这种方法允许对所有大小的环进行系统分析,因为我们考虑七个残基而不是完整环的结构基序。我们重点分析高度重复的循环单词(观察超过 30 次)。我们的研究表明,在 28274 个观察到的单词中,仅 3310 个高度重复出现的结构单词覆盖了 73% 的循环长度。这些结构词的结构变异性较低(平均 RMSd 为 0.85 Å)。正如预期的那样,这些基序中有一半显示出侧翼区域偏好,但有趣的是,三分之二由短环(少于 12 个残基)和长环共享。此外,一半的重复基序表现出显着水平的氨基酸保守性,至少有四个显着位置,并且 87% 的长环包含至少一个这样的单词。我们通过检测统计上过度代表性的结构字母模式来补充我们的分析,就像传统的 DNA 序列分析一样。大约 30% (930) 的结构词被过度表示,并覆盖大约 40% 的循环长度。有趣的是,这些词表现出较低的结构变异性和较高的序列特异性,表明结构或功能的限制。我们开发了一种使用重复结构基序系统地分解和研究蛋白质环的方法。该方法基于结构字母表 HMM-SA,而不是结构对齐和几何参数。我们提取了在短环和长环中发现的有意义的结构图案。据我们所知,这是模式挖掘首次有助于提高蛋白质环中的信噪比。这一发现有助于更好地描述蛋白质环,并可能降低长环分析的复杂性。详细结果请参见http://www.mti.univ-paris-diderot.fr/publication/supplementary/2009/ACCLoop/。
Protein loops encompass 50% of protein residues in available three-dimensional structures. These regions are often involved in protein functions, e.g. binding site, catalytic pocket... However, the description of protein loops with conventional tools is an uneasy task. Regular secondary structures, helices and strands, have been widely studied whereas loops, because they are highly variable in terms of sequence and structure, are difficult to analyze. Due to data sparsity, long loops have rarely been systematically studied. We developed a simple and accurate method that allows the description and analysis of the structures of short and long loops using structural motifs without restriction on loop length. This method is based on the structural alphabet HMM-SA. HMM-SA allows the simplification of a three-dimensional protein structure into a one-dimensional string of states, where each state is a four-residue prototype fragment, called structural letter. The difficult task of the structural grouping of huge data sets is thus easily accomplished by handling structural letter strings as in conventional protein sequence analysis. We systematically extracted all seven-residue fragments in a bank of 93000 protein loops and grouped them according to the structural-letter sequence, named structural word. This approach permits a systematic analysis of loops of all sizes since we consider the structural motifs of seven residues rather than complete loops. We focused the analysis on highly recurrent words of loops (observed more than 30 times). Our study reveals that 73% of loop-lengths are covered by only 3310 highly recurrent structural words out of 28274 observed words). These structural words have low structural variability (mean RMSd of 0.85 Å). As expected, half of these motifs display a flanking-region preference but interestingly, two thirds are shared by short (less than 12 residues) and long loops. Moreover, half of recurrent motifs exhibit a significant level of amino-acid conservation with at least four significant positions and 87% of long loops contain at least one such word. We complement our analysis with the detection of statistically over-represented patterns of structural letters as in conventional DNA sequence analysis. About 30% (930) of structural words are over-represented, and cover about 40% of loop lengths. Interestingly, these words exhibit lower structural variability and higher sequential specificity, suggesting structural or functional constraints. We developed a method to systematically decompose and study protein loops using recurrent structural motifs. This method is based on the structural alphabet HMM-SA and not on structural alignment and geometrical parameters. We extracted meaningful structural motifs that are found in both short and long loops. To our knowledge, it is the first time that pattern mining helps to increase the signal-to-noise ratio in protein loops. This finding helps to better describe protein loops and might permit to decrease the complexity of long-loop analysis. Detailed results are available at http://www.mti.univ-paris-diderot.fr/publication/supplementary/2009/ACCLoop/.
DOI: 10.1016/0014-5793(91)80706-9
发表时间: 1991-06-24
期刊: FEBS LETTERS
影响因子: 3.5
作者:
EFIMOV, AV
通讯作者: EFIMOV, AV
DOI: 10.1093/bioinformatics/btl382
发表时间: 2006-09-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Espadaler, Jordi;Querol, Enrique;Oliva, Baldo
通讯作者: Oliva, Baldo
DOI: 10.1002/prot.10309
发表时间: 2003-03-01
期刊: PROTEINS-STRUCTURE FUNCTION AND GENETICS
影响因子: --
作者:
Hunter, CG;Subramaniam, S
通讯作者: Subramaniam, S
DOI: 10.1186/1471-2105-9-312
发表时间: 2008-07-17
期刊: BMC bioinformatics
影响因子: 3
作者:
Golovin A;Henrick K
通讯作者: Henrick K
DOI: 10.1016/0022-2836(91)80075-6
发表时间: 1991-09-20
影响因子: 5.6
作者:
COLLOCH, N;COHEN, FE
通讯作者: COHEN, FE