KmerStream: streaming algorithms for k-mer abundance estimation

KmerStream: streaming algorithms for k-mer abundance estimation
复制标题

DOI:
10.1093/bioinformatics/btu713
复制
发表时间:
2014-12-15
期刊:
影响因子:
5.8
通讯作者:
Halldorsson, Bjarni V.
Halldorsson, Bjarni V.
中科院分区:
生物学3区
文献类型:
--
作者:
Melsted, Pall;Halldorsson, Bjarni V.

文献摘要

被引文献

相似文献

动机:生物信息学中的一些应用,如基因组组装和纠错方法,依赖于计数和跟踪k-mers(长度为k的子串)。直方图的k-mer频率可以提供有价值的深入了解潜在的分布,并表明在测序experiment.Results中采样的错误率和基因组大小:我们提出KmerStream,流算法估计不同的k-mer存在于高通量测序数据的数量。该算法在时间上与输入的大小成线性关系,而空间需求与输入的大小成对数关系。我们推导出一个简单的模型,使我们能够估计测序实验的错误率,以及基因组大小,仅使用KmerStream报告的聚合统计作为一个应用程序,我们展示了如何KmerStream可以用来计算DNA测序实验的错误率。我们在一组2656个全基因组测序个体上运行KmerStream,并将错误率与测序设备报告的质量值进行比较。我们发现,虽然单独的质量值在很大程度上是可靠的错误率的预测,有相当大的变异性之间的测序运行的错误率,即使当占报告的质量值。
Motivation: Several applications in bioinformatics, such as genome assemblers and error corrections methods, rely on counting and keeping track of k-mers ( substrings of length k). Histograms of k-mer frequencies can give valuable insight into the underlying distribution and indicate the error rate and genome size sampled in the sequencing experiment.Results: We present KmerStream, a streaming algorithm for estimating the number of distinct k-mers present in high-throughput sequencing data. The algorithm runs in time linear in the size of the input and the space requirement are logarithmic in the size of the input. We derive a simple model that allows us to estimate the error rate of the sequencing experiment, as well as the genome size, using only the aggregate statistics reported by KmerStream.As an application we show how KmerStream can be used to compute the error rate of a DNA sequencing experiment. We run KmerStream on a set of 2656 whole genome sequenced individuals and compare the error rate to quality values reported by the sequencing equipment. We discover that while the quality values alone are largely reliable as a predictor of error rate, there is considerable variability in the error rates between sequencing runs, even when accounting for reported quality values.