Calculating orthologs in bacteria and Archaea: a divide and conquer approach.
Calculating orthologs in bacteria and Archaea: a divide and conquer approach.
复制标题
DOI:
10.1371/journal.pone.0028388
复制
发表时间:
2011
期刊:
影响因子:
3.7
通讯作者:
Pallen MJ
中科院分区:
文献类型:
--
作者:
Halachev MR;Loman NJ;Pallen MJ
Among proteins, orthologs are defined as those that are derived by vertical descent from a single progenitor in the last common ancestor of their host organisms. Our goal is to compute a complete set of protein orthologs derived from all currently available complete bacterial and archaeal genomes. Traditional approaches typically rely on all-against-all BLAST searching which is prohibitively expensive in terms of hardware requirements or computational time (requiring an estimated 18 months or more on a typical server). Here, we present xBASE-Orth, a system for ongoing ortholog annotation, which applies a “divide and conquer” approach and adopts a pragmatic scheme that trades accuracy for speed. Starting at species level, xBASE-Orth carefully constructs and uses pan-genomes as proxies for the full collections of coding sequences at each level as it progressively climbs the taxonomic tree using the previously computed data. This leads to a significant decrease in the number of alignments that need to be performed, which translates into faster computation, making ortholog computation possible on a global scale. Using xBASE-Orth, we analyzed an NCBI collection of 1,288 bacterial and 94 archaeal complete genomes with more than 4 million coding sequences in 5 weeks and predicted more than 700 million ortholog pairs, clustered in 175,531 orthologous groups. We have also identified sets of highly conserved bacterial and archaeal orthologs and in so doing have highlighted anomalies in genome annotation and in the proposed composition of the minimal bacterial genome. In summary, our approach allows for scalable and efficient computation of the bacterial and archaeal ortholog annotations. In addition, due to its hierarchical nature, it is suitable for incorporating novel complete genomes and alternative genome annotations. The computed ortholog data and a continuously evolving set of applications based on it are integrated in the xBASE database, available at http://www.xbase.ac.uk/.
登录
查看更多内容
影响因子:
12.3
作者:
Hogg JS;Hu FZ;Janto B;Boissy R;Hayes J;Keefe R;Post JC;Ehrlich GD
通讯作者:
Ehrlich GD
影响因子:
5.5
作者:
Makarova KS;Sorokin AV;Novichkov PS;Wolf YI;Koonin EV
通讯作者:
Koonin EV
影响因子:
4.3
作者:
Altenhoff, Adrian M.;Dessimoz, Christophe
通讯作者:
Dessimoz, Christophe
影响因子:
14.9
作者:
Berglund AC;Sjölund E;Ostlund G;Sonnhammer EL
通讯作者:
Sonnhammer EL
DOI:
10.1142/s0219720008003540
发表时间:
2008-06-01
影响因子:
1
作者:
Fu, Zheng;Jiang, Tao
通讯作者:
Jiang, Tao