An Automated Phylogenetic Tree-Based Small Subunit rRNA Taxonomy and Alignment Pipeline (STAP)

An Automated Phylogenetic Tree-Based Small Subunit rRNA Taxonomy and Alignment Pipeline (STAP)
复制标题

DOI:
10.1371/journal.pone.0002566
复制
发表时间:
2008-07-02
期刊:
影响因子:
3.7
通讯作者:
Eisen, Jonathan A.
Eisen, Jonathan A.
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Wu, Dongying;Hartman, Amber;Eisen, Jonathan A.

文献摘要

被引文献

相似文献

小亚基核糖体RNA(ss-rRNA)基因序列的比较分析形成了我们所知道的培养和未培养微生物的系统发育多样性的基础。随着测序成本的持续下降和通量的增加,ss-rRNA基因的序列正以不断增加的速度获得。这种不断增加的数据流为微生物多样性和进化打开了许多新的窗口,同时也带来了重大的方法学挑战。那些通常需要耗时的人工干预的过程,例如多序列比对的准备,根本无法跟上输入数据的洪流。需要完全自动化的分析方法。值得注意的是,现有的自动化方法避免了一个或多个步骤,虽然计算成本高或困难,我们认为是重要的。特别是,我们认为建设多序列比对和高质量的系统发育分析的性能是必要的。我们在这里描述我们的全自动ss-rRNA分类和比对管道(STAP)。它生成高质量的多重序列比对和系统发育树,因此可以用于多种目的,包括基于遗传学的分类分配和环境样品中物种多样性的分析。该管道将公开可用的软件包(PHYML,BLASTALW和CLUSTALW)与我们的自动对齐,掩蔽和树解析程序相结合。最重要的是,这种自动化过程产生的结果与手动分析可实现的结果相当,但提供了手动工作无法实现的速度和容量。
Comparative analysis of small-subunit ribosomal RNA (ss-rRNA) gene sequences forms the basis for much of what we know about the phylogenetic diversity of both cultured and uncultured microorganisms. As sequencing costs continue to decline and throughput increases, sequences of ss-rRNA genes are being obtained at an ever-increasing rate. This increasing flow of data has opened many new windows into microbial diversity and evolution, and at the same time has created significant methodological challenges. Those processes which commonly require time-consuming human intervention, such as the preparation of multiple sequence alignments, simply cannot keep up with the flood of incoming data. Fully automated methods of analysis are needed. Notably, existing automated methods avoid one or more steps that, though computationally costly or difficult, we consider to be important. In particular, we regard both the building of multiple sequence alignments and the performance of high quality phylogenetic analysis to be necessary. We describe here our fully-automated ss-rRNA taxonomy and alignment pipeline (STAP). It generates both high-quality multiple sequence alignments and phylogenetic trees, and thus can be used for multiple purposes including phylogenetically-based taxonomic assignments and analysis of species diversity in environmental samples. The pipeline combines publicly-available packages (PHYML, BLASTN and CLUSTALW) with our automatic alignment, masking, and tree-parsing programs. Most importantly, this automated process yields results comparable to those achievable by manual analysis, yet offers speed and capacity that are unattainable by manual efforts.