k-SLAM: accurate and ultra-fast taxonomic classification and gene identification for large metagenomic data sets.

k-SLAM: accurate and ultra-fast taxonomic classification and gene identification for large metagenomic data sets.
复制标题

DOI:
10.1093/nar/gkw1248
复制
发表时间:
2017-02-28
影响因子:
14.9
通讯作者:
Butcher SA
Butcher SA
中科院分区:
生物学2区
文献类型:
--
作者:
Ainsworth D;Sternberg MJE;Raczy C;Butcher SA

文献摘要

被引文献

相似文献

k-SLAM是用于宏基因组数据表征的高效算法。与其他超快速宏基因组分类器不同,除了准确的分类分类外,还进行全序列比对,从而允许基因鉴定和变体调用。基于k-mer的方法提供比其他分类器更大的分类准确性,并且比基于比对的方法速度增加三个数量级。使用比对来发现变体和基因沿着它们的分类学起源使得能够表征新菌株。k-SLAM的速度允许在现代大型数据集上进行完整的分类学分类和基因识别。一个伪组装方法是用来提高分类精度高达40%的物种,在其属内具有高序列同源性。
k-SLAM is a highly efficient algorithm for the characterization of metagenomic data. Unlike other ultra-fast metagenomic classifiers, full sequence alignment is performed allowing for gene identification and variant calling in addition to accurate taxonomic classification. A k-mer based method provides greater taxonomic accuracy than other classifiers and a three orders of magnitude speed increase over alignment based approaches. The use of alignments to find variants and genes along with their taxonomic origins enables novel strains to be characterized. k-SLAM's speed allows a full taxonomic classification and gene identification to be tractable on modern large data sets. A pseudo-assembly method is used to increase classification accuracy by up to 40% for species which have high sequence homology within their genus.