Improved software detection and extraction of ITS1 and ITS2 from ribosomal ITS sequences of fungi and other eukaryotes for analysis of environmental sequencing data

Improved software detection and extraction of ITS1 and ITS2 from ribosomal ITS sequences of fungi and other eukaryotes for analysis of environmental sequencing data
复制标题

DOI:
10.1111/2041-210x.12073
复制
发表时间:
2013-10-01
影响因子:
6.6
通讯作者:
Nilsson, R. Henrik
Nilsson, R. Henrik
中科院分区:
环境科学与生态学1区
文献类型:
--
作者:
Bengtsson-Palme, Johan;Ryberg, Martin;Nilsson, R. Henrik

文献摘要

被引文献

相似文献

核内核糖体转录间隔区(ITS)是真菌分子鉴定的主要选择。它的两个高度可变的间隔区(ITS 1和ITS 2)通常是物种特异性的,而插入的5.8S基因是高度保守的。对于序列聚类和blast搜索,依赖于可变间隔区中的任一个而不是保守的5.8S基因通常是有利的。然而,从大型分类学和环境数据集中识别和提取ITS 1和ITS 2通常很困难,并且公共序列数据库中的许多ITS序列界限不正确。我们引入ITSx,这是一个基于Perl的软件工具,用于提取ITS 1、5.8S和ITS 2-以及全长ITS序列-从桑格和高通量测序数据集中。ITSx使用隐马尔可夫模型计算的大型比对共20组真核生物,包括真菌,后生动物和植物,和序列提取是基于预测的核糖体基因的序列中的位置。ITSx具有非常高比例的真阳性提取和低比例的假阳性提取。另外,过程并行化允许非常大的数据集的有利分析,例如一百万个序列扩增子焦磷酸测序数据集。ITSx具有丰富的功能,并且可以轻松集成到自动化序列分析管道中。ITSx为真核生物ITS区域的更灵敏的blast搜索和序列聚类操作铺平了道路。该软件还允许从任何数据集中消除非ITS序列。这对于基于扩增子的下一代测序数据集特别有用,其中在靶序列中经常发现潜在的非靶序列。这样的非靶序列难以通过其他手段找到,并且如果留在数据集中,则会对多样性估计产生噪声。
The nuclear ribosomal internal transcribed spacer (ITS) region is the primary choice for molecular identification of fungi. Its two highly variable spacers (ITS1 and ITS2) are usually species specific, whereas the intercalary 5.8S gene is highly conserved. For sequence clustering and blast searches, it is often advantageous to rely on either one of the variable spacers but not the conserved 5.8S gene. To identify and extract ITS1 and ITS2 from large taxonomic and environmental data sets is, however, often difficult, and many ITS sequences are incorrectly delimited in the public sequence databases.We introduce ITSx, a Perl-based software tool to extract ITS1, 5.8S and ITS2 - as well as full-length ITS sequences - from both Sanger and high-throughput sequencing data sets. ITSx uses hidden Markov models computed from large alignments of a total of 20 groups of eukaryotes, including fungi, metazoans and plants, and the sequence extraction is based on the predicted positions of the ribosomal genes in the sequences. ITSx has a very high proportion of true-positive extractions and a low proportion of false-positive extractions. Additionally, process parallelization permits expedient analyses of very large data sets, such as a one million sequence amplicon pyrosequencing data set. ITSx is rich in features and written to be easily incorporated into automated sequence analysis pipelines.ITSx paves the way for more sensitive blast searches and sequence clustering operations for the ITS region in eukaryotes. The software also permits elimination of non-ITS sequences from any data set. This is particularly useful for amplicon-based next-generation sequencing data sets, where insidious non-target sequences are often found among the target sequences. Such non-target sequences are difficult to find by other means and would contribute noise to diversity estimates if left in the data set.