Calculating orthologs in bacteria and Archaea: a divide and conquer approach.

Calculating orthologs in bacteria and Archaea: a divide and conquer approach.
复制标题

DOI:
10.1371/journal.pone.0028388
复制
发表时间:
2011
期刊:
影响因子:
3.7
通讯作者:
Pallen MJ
Pallen MJ
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Halachev MR;Loman NJ;Pallen MJ

文献摘要

参考文献

被引文献

相似文献

在蛋白质中,直向同源物被定义为通过垂直血统从其宿主生物体的最后共同祖先中的单个祖先衍生的那些。我们的目标是计算一套完整的蛋白质直向同源物来自所有目前可用的完整的细菌和古细菌基因组。传统的方法通常依赖于全对全BLAST搜索,这在硬件要求或计算时间方面非常昂贵(在典型的服务器上需要估计18个月或更长时间)。在这里,我们提出了xBASE-Orth,一个系统正在进行的ortholog注释,它采用了“分而治之”的方法,并采用了务实的计划,交易的准确性和速度。从物种水平开始,xBASE-Orth仔细构建并使用泛基因组作为每个水平的编码序列完整集合的代理,因为它使用先前计算的数据逐步爬上分类树。这导致需要执行的比对数量显著减少,这转化为更快的计算,使得在全球范围内进行直系同源物计算成为可能。使用xBASE-Orth,我们在5周内分析了NCBI收集的1,288个细菌和94个古细菌完整基因组,其中编码序列超过400万,并预测了超过7亿个直系同源物对,聚集在175,531个直系同源物组中。我们还确定了一组高度保守的细菌和古细菌的直系同源物,并在这样做突出了基因组注释和最小细菌基因组的拟议组成中的异常。总之,我们的方法允许细菌和古细菌直系同源物注释的可扩展和有效的计算。此外,由于其层次性,它适合于纳入新的完整基因组和替代基因组注释。计算的直系同源物数据和基于它的一组不断发展的应用程序集成在xBASE数据库中,可在http://www.xbase.ac.uk/上获得。
Among proteins, orthologs are defined as those that are derived by vertical descent from a single progenitor in the last common ancestor of their host organisms. Our goal is to compute a complete set of protein orthologs derived from all currently available complete bacterial and archaeal genomes. Traditional approaches typically rely on all-against-all BLAST searching which is prohibitively expensive in terms of hardware requirements or computational time (requiring an estimated 18 months or more on a typical server). Here, we present xBASE-Orth, a system for ongoing ortholog annotation, which applies a “divide and conquer” approach and adopts a pragmatic scheme that trades accuracy for speed. Starting at species level, xBASE-Orth carefully constructs and uses pan-genomes as proxies for the full collections of coding sequences at each level as it progressively climbs the taxonomic tree using the previously computed data. This leads to a significant decrease in the number of alignments that need to be performed, which translates into faster computation, making ortholog computation possible on a global scale. Using xBASE-Orth, we analyzed an NCBI collection of 1,288 bacterial and 94 archaeal complete genomes with more than 4 million coding sequences in 5 weeks and predicted more than 700 million ortholog pairs, clustered in 175,531 orthologous groups. We have also identified sets of highly conserved bacterial and archaeal orthologs and in so doing have highlighted anomalies in genome annotation and in the proposed composition of the minimal bacterial genome. In summary, our approach allows for scalable and efficient computation of the bacterial and archaeal ortholog annotations. In addition, due to its hierarchical nature, it is suitable for incorporating novel complete genomes and alternative genome annotations. The computed ortholog data and a continuously evolving set of applications based on it are integrated in the xBASE database, available at http://www.xbase.ac.uk/.
DOI: 10.1186/gb-2007-8-6-r103
发表时间: 2007
期刊: Genome biology
影响因子: 12.3
作者:
Hogg JS;Hu FZ;Janto B;Boissy R;Hayes J;Keefe R;Post JC;Ehrlich GD
通讯作者: Ehrlich GD
DOI: 10.1186/1745-6150-2-33
发表时间: 2007-11-27
期刊: Biology direct
影响因子: 5.5
作者:
Makarova KS;Sorokin AV;Novichkov PS;Wolf YI;Koonin EV
通讯作者: Koonin EV
DOI: 10.1371/journal.pcbi.1000262
发表时间: 2009-01
影响因子: 4.3
作者:
Altenhoff, Adrian M.;Dessimoz, Christophe
通讯作者: Dessimoz, Christophe
非par体6:具有内核的真核直系同源群。
DOI: 10.1093/nar/gkm1020
发表时间: 2008-01
影响因子: 14.9
作者:
Berglund AC;Sjölund E;Ostlund G;Sonnhammer EL
通讯作者: Sonnhammer EL
DOI: 10.1142/s0219720008003540
发表时间: 2008-06-01
影响因子: 1
作者:
Fu, Zheng;Jiang, Tao
通讯作者: Jiang, Tao