BASTA - Taxonomic classification of sequences and sequence bins using last common ancestor estimations

BASTA - Taxonomic classification of sequences and sequence bins using last common ancestor estimations
复制标题

DOI:
10.1111/2041-210x.13095
复制
发表时间:
2019-01-01
影响因子:
6.6
通讯作者:
Ralph, Peter J.
Ralph, Peter J.
中科院分区:
环境科学与生态学1区
文献类型:
--
作者:
Kahlke, Tim;Ralph, Peter J.

文献摘要

被引文献

相似文献

鉴定DNA序列的分类起源对许多测序项目至关重要,例如宏基因组学研究、全基因组测序项目中的污染物鉴定以及基于标记基因的群落分析中感兴趣的生物筛选。最后共同祖先算法是估计给定序列分类的有效方法,已广泛用于下一代测序(NGS) reads的分类,也称为第二代测序reads。在这里,我们提出了一个基本序列分类注释器BASTA(),它利用用户定义的最佳匹配的NCBI分类,将最后共同祖先估计从测序读取扩展到任何类型的核苷酸或氨基酸序列。BASTA可以配置为使用许多常见序列比较工具的输出,例如BLAST和Diamond,并与提供的或用户定义的目标序列数据库结合使用。
Identification of the taxonomic origin of a DNA sequence is crucial for many sequencing projects, e.g. metagenomics studies, identification of contaminations in whole genome sequencing projects and filtering of organisms of interest in marker-gene based community analyses. Last common ancestor algorithms are powerful approaches to estimate the taxonomy of a given sequence and have been widely used for classification of next-generation sequencing (NGS) reads, also known as 2nd generation sequencing reads. Here, we present BASTA (), a basic sequence taxonomy annotator, which extends last common ancestor estimations from sequencing reads to any kind of nucleotide or amino acid sequence utilizing NCBI taxonomies of user-defined best hits. BASTA can be configured to use the output of many common sequence comparison tools, e.g. BLAST and Diamond, in conjunction with either provided or user-defined target sequence databases.