Phylogenetic and functional assessment of orthologs inference projects and methods.

Phylogenetic and functional assessment of orthologs inference projects and methods.
复制标题

DOI:
10.1371/journal.pcbi.1000262
复制
发表时间:
2009-01
影响因子:
4.3
通讯作者:
Dessimoz, Christophe
Dessimoz, Christophe
中科院分区:
生物学2区
文献类型:
--
作者:
Altenhoff, Adrian M.;Dessimoz, Christophe

文献摘要

参考文献

被引文献

相似文献

在全基因组范围内准确鉴定同源基因是比较基因组学的一个核心问题,这一事实反映在近年来发展的许多同源鉴定项目中。然而,只有少数报告比较了它们的准确性,事实上,最近的几项努力尚未得到系统的评价。此外,尽管惠誉最初的定义是基于遗传学的,但通常只从功能守恒的角度来评估正形学。我们收集并绘制了九个领先的矫正项目和方法(COG,KOG,Inparanoid,OrthoMCL,Ensembl Compara,Homologene,RoundUp,EggNOG和OMA)和两个标准方法(双向最佳命中和倒数最小距离)的结果。我们使用六种不同的测试系统地比较了他们在胚胎发生和功能方面的预测。这需要映射数百万个序列,处理数亿个预测的直系同源物对,以及计算数万棵树。在系统发育分析或功能分析中,需要高特异性,我们发现OMA和同源基因表现最好。在较低的功能特异性但较高的覆盖水平下,OrthoMCL优于Ensembl Compara,并且在较小程度上优于Inparanoid。最后,最近的EggNOG的大覆盖率可以用于建立广泛的功能分组,但该方法对于系统发育或详细的功能分析不够特异。在一般的方法,我们观察到,更复杂的树重建/和解的方法Ensembl比较有时优于成对比较的方法,即使在系统发育测试。此外,我们表明,标准的双向最佳命中往往优于项目更复杂的算法。首先,本研究提供了指导的正字法数据用户的广泛社区,以数据库最适合他们的需求。第二,引入了新的正字法验证方法。第三,它为当前和未来的方法设定了性能标准。直向同源物的鉴定是基因组学的一个核心问题,在许多研究领域都有应用,包括比较基因组学、遗传学、蛋白质功能注释和基因组重排。越来越多的项目旨在从完整的基因组中推断直系同源物,但对其相对准确性或覆盖率知之甚少。由于整个基因组的确切进化历史在很大程度上仍然是未知的,预测只能间接地验证,也就是说,在不同的应用程序的正形学。到目前为止,发表的少数比较研究已经完全从直向同源物具有保守的蛋白质功能的期望中评估了直向同源物。在目前的工作中,我们引入的方法来验证同源性方面的同源性,并进行了全面的比较九个领先的同源推理项目和两种方法,使用系统发育和功能测试。结果表明,不同项目的性能差异很大,这表明,选择的orthology数据库可以有很大的影响,任何下游分析。
Accurate genome-wide identification of orthologs is a central problem in comparative genomics, a fact reflected by the numerous orthology identification projects developed in recent years. However, only a few reports have compared their accuracy, and indeed, several recent efforts have not yet been systematically evaluated. Furthermore, orthology is typically only assessed in terms of function conservation, despite the phylogeny-based original definition of Fitch. We collected and mapped the results of nine leading orthology projects and methods (COG, KOG, Inparanoid, OrthoMCL, Ensembl Compara, Homologene, RoundUp, EggNOG, and OMA) and two standard methods (bidirectional best-hit and reciprocal smallest distance). We systematically compared their predictions with respect to both phylogeny and function, using six different tests. This required the mapping of millions of sequences, the handling of hundreds of millions of predicted pairs of orthologs, and the computation of tens of thousands of trees. In phylogenetic analysis or in functional analysis where high specificity is required, we find that OMA and Homologene perform best. At lower functional specificity but higher coverage level, OrthoMCL outperforms Ensembl Compara, and to a lesser extent Inparanoid. Lastly, the large coverage of the recent EggNOG can be of interest to build broad functional grouping, but the method is not specific enough for phylogenetic or detailed function analyses. In terms of general methodology, we observe that the more sophisticated tree reconstruction/reconciliation approach of Ensembl Compara was at times outperformed by pairwise comparison approaches, even in phylogenetic tests. Furthermore, we show that standard bidirectional best-hit often outperforms projects with more complex algorithms. First, the present study provides guidance for the broad community of orthology data users as to which database best suits their needs. Second, it introduces new methodology to verify orthology. And third, it sets performance standards for current and future approaches. The identification of orthologs, pairs of homologous genes in different species that started diverging through speciation events, is a central problem in genomics with applications in many research areas, including comparative genomics, phylogenetics, protein function annotation, and genome rearrangement. An increasing number of projects aim at inferring orthologs from complete genomes, but little is known about their relative accuracy or coverage. Because the exact evolutionary history of entire genomes remains largely unknown, predictions can only be validated indirectly, that is, in the context of the different applications of orthology. The few comparison studies published so far have asssessed orthology exclusively from the expectation that orthologs have conserved protein function. In the present work, we introduce methodology to verify orthology in terms of phylogeny and perform a comprehensive comparison of nine leading ortholog inference projects and two methods using both phylogenetic and functional tests. The results show large variations among the different projects in terms of performances, which indicates that the choice of orthology database can have a strong impact on any downstream analysis.
非par体6:具有内核的真核直系同源群。
DOI: 10.1093/nar/gkm1020
发表时间: 2008-01
影响因子: 14.9
作者:
Berglund AC;Sjölund E;Ostlund G;Sonnhammer EL
通讯作者: Sonnhammer EL
DOI: 10.1093/nar/gki913
发表时间: 2005
影响因子: 14.9
作者:
Notebaart RA;Huynen MA;Teusink B;Siezen RJ;Snel B
通讯作者: Snel B
DOI: 10.1093/nar/gkl440
发表时间: 2006
影响因子: 14.9
作者:
Bern M;Goldberg D;Lyashenko E
通讯作者: Lyashenko E
DOI: 10.1371/journal.pcbi.0010045
发表时间: 2005-10
影响因子: 4.3
作者:
Engelhardt BE;Jordan MI;Muratore KE;Brenner SE
通讯作者: Brenner SE
DOI: 10.1093/nar/gkh036
发表时间: 2004-01-01
影响因子: 14.9
作者:
Harris, MA;Clark, J;White, R
通讯作者: White, R