Protein molecular function prediction by Bayesian phylogenomics.

Protein molecular function prediction by Bayesian phylogenomics.
复制标题

DOI:
10.1371/journal.pcbi.0010045
复制
发表时间:
2005-10
影响因子:
4.3
通讯作者:
Brenner SE
Brenner SE
中科院分区:
生物学2区
文献类型:
--
作者:
Engelhardt BE;Jordan MI;Muratore KE;Brenner SE

文献摘要

参考文献

被引文献

相似文献

我们提出了一个统计图形模型来推断特定的分子功能,为未注释的蛋白质序列使用同源性。基于系统基因组学原理,SIFTER(通过进化关系的功能统计推断)在给定协调的系统发育和可用功能注释的情况下,即使数据稀疏或有噪声,也能准确预测蛋白质家族成员的分子功能。与Gene Ontology注释数据库、BLAST、GOtcha和Orthostrapper相比,我们的方法在100个Pfam家族中产生了特异性和一致的分子功能预测。我们对腺苷-5 ' -单磷酸腺苷/腺苷脱氨酶家族和乳酸/苹果酸脱氢酶家族的功能预测进行了更详细的探索,在前者的情况下,将预测与已发表的功能表征的金标准集进行了比较。考虑到脱氨酶家族中3%的蛋白质的功能注释,SIFTER在预测实验表征蛋白质的分子功能方面达到了96%的准确性。与BLAST(75%)、GeneQuiz(64%)、GOtcha(89%)和Orthostrapper(11%)等现有方法相比,SIFTER在该数据集上的准确性有了显著提高。我们还通过实验表征了恶性疟原虫的腺苷脱氨酶,证实了SIFTER的预测。这些结果说明了在系统基因组学问题中利用功能进化的统计模型的预测能力。SIFTER的软件实现可从作者处获得。新的基因组序列继续以惊人的速度发表。然而,未注释的序列对生物学家的使用是有限的。为了对假设的蛋白质进行分子功能的计算注释,研究人员通常试图从进化相关的蛋白质中进行某种形式的信息传递。这种转移是在系统发育关系的背景下最成功地实现的,利用了在给定蛋白质家族中关于分子进化的全面知识。这种分子功能注释的通用方法被称为系统基因组学,它是目前提供高质量注释的最佳方法。然而,系统基因组学的一个缺点是,它是一个耗时的手动过程,需要专业知识。在当前的论文中,作者已经开发了一种统计方法——被称为SIFTER(通过进化关系的功能统计推断)——它允许系统基因组分析自动进行。作者介绍了在100个蛋白质家族集合上运行SIFTER的结果。他们还在一个特定的家族上验证他们的方法,该家族有一组金标准的实验注释可用。他们表明,SIFTER正确注释了96%的金标准蛋白,优于流行的注释方法,包括基于blast的注释(75%)、GOtcha(89%)、GeneQuiz(64%)和Orthostrapper(11%)。该结果支持了开展高质量全基因组系统基因组分析的可行性。
We present a statistical graphical model to infer specific molecular function for unannotated protein sequences using homology. Based on phylogenomic principles, SIFTER (Statistical Inference of Function Through Evolutionary Relationships) accurately predicts molecular function for members of a protein family given a reconciled phylogeny and available function annotations, even when the data are sparse or noisy. Our method produced specific and consistent molecular function predictions across 100 Pfam families in comparison to the Gene Ontology annotation database, BLAST, GOtcha, and Orthostrapper. We performed a more detailed exploration of functional predictions on the adenosine-5′-monophosphate/adenosine deaminase family and the lactate/malate dehydrogenase family, in the former case comparing the predictions against a gold standard set of published functional characterizations. Given function annotations for 3% of the proteins in the deaminase family, SIFTER achieves 96% accuracy in predicting molecular function for experimentally characterized proteins as reported in the literature. The accuracy of SIFTER on this dataset is a significant improvement over other currently available methods such as BLAST (75%), GeneQuiz (64%), GOtcha (89%), and Orthostrapper (11%). We also experimentally characterized the adenosine deaminase from Plasmodium falciparum, confirming SIFTER's prediction. The results illustrate the predictive power of exploiting a statistical model of function evolution in phylogenomic problems. A software implementation of SIFTER is available from the authors. New genome sequences continue to be published at a prodigious rate. However, unannotated sequences are of limited use to biologists. To computationally annotate a hypothetical protein for molecular function, researchers generally attempt to carry out some form of information transfer from evolutionarily related proteins. Such transfer is most successfully achieved within the context of phylogenetic relationships, exploiting the comprehensive knowledge that is available regarding molecular evolution within a given protein family. This general approach to molecular function annotation is known as phylogenomics, and it is the best method currently available for providing high-quality annotations. A drawback of phylogenomics, however, is that it is a time-consuming manual process requiring expert knowledge. In the current paper, the authors have developed a statistical approach—referred to as SIFTER (Statistical Inference of Function Through Evolutionary Relationships)—that allows phylogenomic analyses to be carried out automatically. The authors present the results of running SIFTER on a collection of 100 protein families. They also validate their method on a specific family for which a gold standard set of experimental annotations is available. They show that SIFTER annotates 96% of the gold standard proteins correctly, outperforming popular annotation methods including BLAST-based annotation (75%), GOtcha (89%), GeneQuiz (64%), and Orthostrapper (11%). The results support the feasibility of carrying out high-quality phylogenomic analyses of entire genomes.
DOI: 10.1093/nar/gkg005
发表时间: 2003-01-01
影响因子: 14.9
作者:
Frishman, D;Mokrejs, M;Mewes, HW
通讯作者: Mewes, HW
DOI: 10.1093/bioinformatics/14.7.600
发表时间: 1998-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Andrade, MA;Valencia, A
通讯作者: Valencia, A
DOI: 10.2307/2412448
发表时间: 1970-01-01
期刊: SYSTEMATIC ZOOLOGY
影响因子: --
作者:
FITCH, WM
通讯作者: FITCH, WM
DOI: 10.1093/nar/gkp985
发表时间: 2010-01
影响因子: 14.9
作者:
Finn RD;Mistry J;Tate J;Coggill P;Heger A;Pollington JE;Gavin OL;Gunasekaran P;Ceric G;Forslund K;Holm L;Sonnhammer EL;Eddy SR;Bateman A
通讯作者: Bateman A
DOI: 10.1016/0168-9525(96)81406-5
发表时间: 1996-02-01
期刊: TRENDS IN GENETICS
影响因子: 11.4
作者:
Gaasterland, T;Sensen, CW
通讯作者: Sensen, CW