The Bologna Annotation Resource: a Non Hierarchical Method for the Functional and Structural Annotation of Protein Sequences Relying on a Comparative Large-Scale Genome Analysis

The Bologna Annotation Resource: a Non Hierarchical Method for the Functional and Structural Annotation of Protein Sequences Relying on a Comparative Large-Scale Genome Analysis
复制标题

DOI:
10.1021/pr900204r
复制
发表时间:
2009-09-01
影响因子:
4.4
通讯作者:
Casadio, Rita
Casadio, Rita
中科院分区:
生物学2区
文献类型:
--
作者:
Bartoli, Lisa;Montanucci, Ludovica;Casadio, Rita

文献摘要

被引文献

相似文献

蛋白质序列注释是后基因组时代的一个重大挑战。由于完整的基因组和蛋白质组的可用性,蛋白质注释最近从跨基因组比较中获得了宝贵的优势。在这项工作中,我们描述了一个新的非层次聚类程序,其特征在于一个严格的度量,确保可靠的转移功能相关的蛋白质之间,即使在多域和远亲的蛋白质的情况下。该方法利用599完全测序的基因组,无论是从原核生物和真核生物,和GO和PDB/SCOP映射的集群的比较分析。我们的方法的统计验证表明,我们的聚类技术捕获同源和远亲蛋白质序列之间共享的基本信息。通过这样,可以通过继承聚类的注释来安全地注释未表征的蛋白质。我们通过盲目注释其他201个基因组来验证我们的方法,最后我们开发了BAR(博洛尼亚注释资源),这是一个基于总共800个基因组的蛋白质功能注释的预测服务器(可在http://microserf.biocomp.unibo.it/bar/上公开获得)。
Protein sequence annotation is a major challenge in the postgenomic era. Thanks to the availability of complete genomes and proteomes, protein annotation has recently taken invaluable advantage from cross-genome comparisons. In this work, we describe a new non hierarchical clustering procedure characterized by a stringent metric which ensures a reliable transfer of function between related proteins even in the case of multidomain and distantly related proteins. The method takes advantage of the comparative analysis of 599 completely sequenced genomes, both from prokaryotes and eukaryotes, and of a GO and PDB/SCOP mapping over the clusters. A statistical validation of our method demonstrates that our clustering technique captures the essential information shared between homologous and distantly related protein sequences. By this, uncharacterized proteins can be safely annotated by inheriting the annotation of the cluster. We validate our method by blindly annotating other 201 genomes and finally we develop BAR (the Bologna Annotation Resource), a prediction server for protein functional annotation based on a total of 800 genomes (publicly available at http://microserf.biocomp.unibo.it/bar/).