Comprehensive aligned sequence construction for automated design of effective probes (CASCADE-P) using 16S rDNA

Comprehensive aligned sequence construction for automated design of effective probes (CASCADE-P) using 16S rDNA
复制标题

DOI:
10.1093/bioinformatics/btg200
复制
发表时间:
2003-08-12
期刊:
影响因子:
5.8
通讯作者:
Andersen, GL
Andersen, GL
中科院分区:
生物学3区
文献类型:
--
作者:
DeSantis, TZ;Dubosarskiy, I;Andersen, GL

文献摘要

被引文献

相似文献

动机:已利用16S rRNA基因的序列变异来鉴定原核生物。这些变异引导用于检测分类群或特定生物的DNA探针的设计。我们项目的长期目标是创建能够识别未知样本中16S rDNA序列的探针阵列。这就需要对超过75000个公开可用的“16S”序列进行验证、分类和比对。理想情况下,整个过程应该通过计算机管理,以便比对后的集合能够定期吸收公共记录中的16S rDNA序列。一个完整的多序列比对将为计算机选择探针提供基础,并有助于微生物分类学和系统发育研究。 结果:在此我们报道了62662个16S rDNA序列的比对和相似性聚类,以及一种为每个聚类设计有效探针的方法。设计了一种新的比对压缩算法——NAST(最近比对空间终止),以产生被称为原核生物多序列比对(prokMSA)的统一多序列比对。从原核生物多序列比对中,基于传递序列相似性发现了9020个操作分类单元(OTUs)。利用聚类为操作分类单元的原核生物多序列比对,一种自动的探针设计方法是直接的。作为一个测试案例,通过计算机为葡萄球菌属内鉴定出的27个操作分类单元中的每一个挑选了多个探针。这些探针被整合到一个定制的微阵列中,并且能够将金黄色葡萄球菌和炭疽芽孢杆菌正确分类到它们正确的操作分类单元中。尽管概述了一种成功的探针挑选策略,但创建原核生物多序列比对的主要重点是提供一个全面的、分类的、可更新的16S rDNA集合,作为任何探针选择算法的基础。
Motivation: Prokaryotic organisms have been identified utilizing the sequence variation of the 16S rRNA gene. Variations steer the design of DNA probes for the detection of taxonomic groups or specific organisms. The long-term goal of our project is to create probe arrays capable of identifying 16S rDNA sequences in unknown samples. This necessitated the authentication, categorization and alignment of the >75000 publicly available '16S' sequences. Preferably, the entire process should be computationally administrated so the aligned collection could periodically absorb 16S rDNA sequences from the public records. A complete multiple sequence alignment would provide a foundation for computational probe selection and facilitates microbial taxonomy and phylogeny.Results: Here we report the alignment and similarity clustering of 62 662 16S rDNA sequences and an approach for designing effective probes for each cluster. A novel alignment compression algorithm, NAST (Nearest Alignment Space Termination), was designed to produce the uniform multiple sequence alignment referred to as the prokMSA. From the prokMSA, 9020 Operational Taxonomic Units (OTUs) were found based on transitive sequence similarities. An automated approach to probe design was straightforward using the prokMSA clustered into OTUs. As a test case, multiple probes were computationally picked for each of the 27 OTUs that were identified within the Staphylococcus Group. The probes were incorporated into a customized microarray and were able to correctly categorize Staphylococcus aureus and Bacillus anthracis into their correct OTUs. Although a successful probe picking strategy is outlined, the main focus of creating the prokMSA was to provide a comprehensive, categorized, updateable 16S rDNA collection useful as a foundation for any probe selection algorithm.