SINA: accurate high-throughput multiple sequence alignment of ribosomal RNA genes.

SINA: accurate high-throughput multiple sequence alignment of ribosomal RNA genes.
复制标题

DOI:
10.1093/bioinformatics/bts252
复制
发表时间:
2012-07-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Glöckner FO
Glöckner FO
中科院分区:
其他
文献类型:
--
作者:
Pruesse E;Peplies J;Glöckner FO

文献摘要

参考文献

被引文献

相似文献

动机:在同源序列分析中,多重序列比对(MSA)的计算已成为瓶颈。这对于像核糖体 RNA (rRNA) 这样的标记基因来说尤其麻烦,因为核糖体 RNA (rRNA) 已经有数百万个序列可供公开使用,个别研究可以轻松产生数十万个新序列。已经开发了一些方法来应对这些数字,但需要进一步改进才能满足精度要求。结果:在本研究中,我们提出了 SILVA 增量对准器 (SINA),用于比对 SILVA 核糖体 RNA 项目提供的 rRNA 基因数据库。 SINA结合使用k聚体搜索和偏序比对(POA)来保持非常高的比对精度,同时满足高吞吐量性能需求。 SINA 与常用的高吞吐量 MSA 程序 PyNAST 和 mothur 进行了比较评估。三个 BRAliBase III 基准 MSA 的重现精度分别为 99.3、97.6 和 96.1。使用包含 1000 和 5000 个序列的参考 MSA,可以以 98.9% 和 99.3% 的准确度重现包含 38 772 个序列的较大基准 MSA。在所有执行的基准测试中,SINA 都能获得比 PyNAST 和 mothur 更高的准确度。可用性:http://www.arb-silva.de/aligner 提供使用最新 SILVA SSU/LSU 参考数据集作为参考 MSA 的最多 500 个序列的比对。此页面还链接到 Linux 二进制文件、用户手册和教程。 SINA 是根据个人使用许可提供的。联系方式:epruesse@mpi-bremen.de 补充信息:补充数据可在生物信息学在线获取。
Motivation: In the analysis of homologous sequences, computation of multiple sequence alignments (MSAs) has become a bottleneck. This is especially troublesome for marker genes like the ribosomal RNA (rRNA) where already millions of sequences are publicly available and individual studies can easily produce hundreds of thousands of new sequences. Methods have been developed to cope with such numbers, but further improvements are needed to meet accuracy requirements. Results: In this study, we present the SILVA Incremental Aligner (SINA) used to align the rRNA gene databases provided by the SILVA ribosomal RNA project. SINA uses a combination of k-mer searching and partial order alignment (POA) to maintain very high alignment accuracy while satisfying high throughput performance demands. SINA was evaluated in comparison with the commonly used high throughput MSA programs PyNAST and mothur. The three BRAliBase III benchmark MSAs could be reproduced with 99.3, 97.6 and 96.1 accuracy. A larger benchmark MSA comprising 38 772 sequences could be reproduced with 98.9 and 99.3% accuracy using reference MSAs comprising 1000 and 5000 sequences. SINA was able to achieve higher accuracy than PyNAST and mothur in all performed benchmarks. Availability: Alignment of up to 500 sequences using the latest SILVA SSU/LSU Ref datasets as reference MSA is offered at http://www.arb-silva.de/aligner. This page also links to Linux binaries, user manual and tutorial. SINA is made available under a personal use license. Contact: epruesse@mpi-bremen.de Supplementary information: Supplementary data are available at Bioinformatics online.
DOI: 10.1186/1748-7188-1-19
发表时间: 2006-10-24
期刊: Algorithms for molecular biology : AMB
影响因子: --
作者:
Wilm A;Mainz I;Steger G
通讯作者: Steger G
DOI: 10.1093/nar/gkh293
发表时间: 2004-02-01
影响因子: 14.9
作者:
Ludwig, W;Strunk, O;Schleifer, KH
通讯作者: Schleifer, KH
DOI: 10.1093/nar/gkm864
发表时间: 2007
影响因子: 14.9
作者:
Pruesse E;Quast C;Knittel K;Fuchs BM;Ludwig W;Peplies J;Glöckner FO
通讯作者: Glöckner FO
DOI: 10.1093/nar/gkp998
发表时间: 2010-01
影响因子: 14.9
作者:
Leinonen R;Akhtar R;Birney E;Bonfield J;Bower L;Corbett M;Cheng Y;Demiralp F;Faruque N;Goodgame N;Gibson R;Hoad G;Hunter C;Jang M;Leonard S;Lin Q;Lopez R;Maguire M;McWilliam H;Plaister S;Radhakrishnan R;Sobhany S;Slater G;Ten Hoopen P;Valentin F;Vaughan R;Zalunin V;Zerbino D;Cochrane G
通讯作者: Cochrane G
DOI: 10.1093/oxfordjournals.molbev.a025779
发表时间: 1997-04-01
影响因子: 10.7
作者:
Morrison, DA;Ellis, JT
通讯作者: Ellis, JT