Clustering protein sequences-structure prediction by transitive homology

Clustering protein sequences-structure prediction by transitive homology
复制标题

DOI:
10.1093/bioinformatics/17.10.935
复制
发表时间:
2001-10-01
期刊:
影响因子:
5.8
通讯作者:
Schrader, R
Schrader, R
中科院分区:
生物学3区
文献类型:
--
作者:
Bolten, E;Schliep, A;Schrader, R

文献摘要

被引文献

相似文献

动机:人们普遍认为,对于两个蛋白质A和B,序列同一性高于某个阈值意味着由于共同的进化祖先而具有结构相似性。由于这只是结构相似性的充分条件,但不是必要条件,所以问题仍然是可以用什么其他标准来鉴定远程同源物。传递性是指从第三个蛋白质B的存在推导出蛋白质A和C之间的结构相似性的概念,使得A和B以及B和C是同源物,如果A和B之间以及B和C之间的序列一致性高于上述阈值,则确定A和B之间以及B和C之间的序列一致性。传递性是否总是成立,传递性是否可以无限扩展,目前还不是很清楚。结果:我们提出了一种基于图的聚类方法,其中传递性起着至关重要的作用。我们使用Smith-Waterman局部比对算法确定了SwissProt数据库中序列的所有配对相似性。这些数据被转换成有向图,其中蛋白质序列构成顶点。如果序列A和B的相似性相对于A的自相似性在固定阈值以上,则从顶点A到顶点B绘制有向边。传递性在聚类过程中很重要,因为使用了中间序列,尽管受到在这些序列上连接的蛋白质之间具有双向定向路径的要求的限制。比对分数尺度的自相似性所隐含的长度相关性似乎是避免多域蛋白导致的聚类错误的有效准则。为了处理产生的大图,我们开发了一个高效的图库。方法包括能够处理多结构域蛋白质的新的基于图的聚类算法和聚类比较算法。蛋白质结构分类(SCOP)被用作我们方法的评估数据集,在检测远程同源基因方面比成对比较提高了24%。
Motivation: It is widely believed that for two proteins A and B a sequence identity above some threshold implies structural similarity due to a common evolutionary ancestor. Since this is only a sufficient, but not a necessary condition for structural similarity, the question remains what other criteria can be used to identify remote homologues.Transitivity refers to the concept of deducing a structural similarity between proteins A and C from the existence of a third protein B, such that A and B as well as B and C are homologues, as ascertained if the sequence identity between A and B as well as that between B and C is above the aforementioned threshold. It is not fully understood if transitivity always holds and whether transitivity can be extended ad infinitum.Results: We developed a graph-based clustering approach, where transitivity plays a crucial role. We determined all pair-wise similarities for the sequences in the SwissProt database using the Smith-Waterman local alignment algorithm. This data was transformed into a directed graph, where protein sequences constitute vertices. A directed edge was drawn from vertex A to vertex B if the sequences A and B showed similarity scaled with respect to the self-similarity of A, above a fixed threshold. Transitivity was important in the clustering process, as intermediate sequences were used, limited though by the requirement of having directed paths in both directions between proteins linked over such sequences. The length dependency-implied by the self-similarity-of the scaling of the alignment scores appears to be an effective criterion to avoid clustering errors due to multi-domain proteins.To deal with the resulting large graphs we have developed an efficient library. Methods include the novel graph-based clustering algorithm capable of handling multi-domain proteins and cluster comparison algorithms. Structural Classification of Proteins (SCOP) was used as an evaluation data set for our method, yielding a 24% improvement over pair-wise comparisons in terms of detecting remote homologues.