Error filtering, pair assembly and error correction for next-generation sequencing reads

Error filtering, pair assembly and error correction for next-generation sequencing reads
复制标题

DOI:
10.1093/bioinformatics/btv401
复制
发表时间:
2015-11-01
期刊:
影响因子:
5.8
通讯作者:
Flyvbjerg, Henrik
Flyvbjerg, Henrik
中科院分区:
生物学3区
文献类型:
--
作者:
Edgar, Robert C.;Flyvbjerg, Henrik

文献摘要

被引文献

相似文献

动机:下一代测序产生大量的数据与错误是很难区分从真正的生物变异时,覆盖率低。结果:我们证明了大幅度减少错误频率,特别是对于高错误率读取,通过三个独立的手段:(i)根据其预期的错误数过滤读段,(ii)组装重叠读段对,以及(iii)对于扩增子读段,通过利用独特的序列丰度来执行纠错。我们还表明,大多数已发表的配对读取组装器计算不正确的后验质量分数。
Motivation: Next-generation sequencing produces vast amounts of data with errors that are difficult to distinguish from true biological variation when coverage is low.Results: We demonstrate large reductions in error frequencies, especially for high-error-rate reads, by three independent means: (i) filtering reads according to their expected number of errors, (ii) assembling overlapping read pairs and (iii) for amplicon reads, by exploiting unique sequence abundances to perform error correction. We also show that most published paired read assemblers calculate incorrect posterior quality scores.