Structural similarity to link sequence space: New potential superfamilies and implications for structural genomics

Structural similarity to link sequence space: New potential superfamilies and implications for structural genomics
复制标题

DOI:
10.1110/ps.3950102
复制
发表时间:
2002-05-01
期刊:
影响因子:
8
通讯作者:
Russell, RB
Russell, RB
中科院分区:
生物学3区
文献类型:
--
作者:
Aloy, P;Oliva, B;Russell, RB

文献摘要

被引文献

相似文献

目前结构生物学的发展速度意味着蛋白质的三维结构可以在蛋白质功能之前被知道。通过日益重要的结构比较来确定同源性的方法。先前的研究表明,基于结构的比对后的序列相似性是同源性和通常功能相似性的最佳鉴别器之一。在这里,我们利用这一观察结果,结合蛋白质结构和序列数据库的合并,来预测遥远的同源关系。我们使用蛋白质结构分类(SCOP)数据库连接SMART和Pfam数据库的序列比对。因此,我们提供了新的路线,不能很容易地构建在没有已知的三维结构。然后,我们扩展了Murzin(1993 b)的方法,对结构比对后发现的序列同一性赋予统计学意义,从而提出不同序列家族之间的最佳联系。我们发现,几个远亲的蛋白质序列家族可以与信心,显示的方法是一种手段,用于推断同源关系,从而可能的功能时,蛋白质的结构已知,但未知的功能。分析还发现了几个新的潜在的超家族,其中相关的比对和重叠的检查揭示了不寻常的结构特征或保守的氨基酸和结合底物的共定位的保守性。我们讨论了结构基因组学的倡议和改进序列比较方法的影响。
The current pace of structural biology now means that protein three-dimensional structure can be known before protein function. making methods for assigning homology via structure comparison of growing importance. Previous research has suggested that sequence similarity after structure-based alignment is one of the best discriminators of homology and often functional similarity. Here, we exploit this observation, together with a merger of protein structure and sequence databases, to predict distant homologous relationships. We use the Structural Classification of Proteins (SCOP) database to link sequence alignments from the SMART and Pfam databases. We thus provide new alignments that could not be constructed easily in the absence of known three-dimensional structures. We then extend the method of Murzin ( 1993b) to assign statistical significance to sequence identities found after structural alignment and thus suggest the best link between diverse sequence families. We find that several distantly related protein sequence families can be linked with confidence, showing the approach to be a means for inferring homologous relationships and thus possible functions when proteins are of known structure but of unknown function. The analysis also finds several new potential superfamilies, where inspection of the associated alignments and superimpositions reveals conservation of unusual structural features or co-location of conserved amino acids and bound substrates. We discuss implications for Structural Genomics initiatives and for improvements to sequence comparison methods.