CARE: context-aware sequencing read error correction

CARE: context-aware sequencing read error correction
复制标题

DOI:
10.1093/bioinformatics/btaa738
复制
发表时间:
2021-04-01
期刊:
影响因子:
5.8
通讯作者:
Schmidt, Bertil
Schmidt, Bertil
中科院分区:
生物学3区
文献类型:
--
作者:
Kallenborn, Felix;Hildebrandt, Andreas;Schmidt, Bertil

文献摘要

被引文献

相似文献

动机:纠错是许多下一代测序(NGS)流水线中的基本预处理步骤,特别是对于从头基因组组装。然而,现有的错误校正方法要么遭受高假阳性率,因为他们打破读成独立的k-mer或不扩展有效地大量的测序reads和复杂的genomes.Results:我们提出了护理-一个基于数据库的可扩展的错误校正算法的Illumina数据使用minhashing的概念。Minhashing允许在大型测序读取集合内进行有效的相似性搜索,这使得能够快速计算高质量的多重比对。通过详细检查相应的比对来校正测序错误。我们的性能评估表明,CARE生成的假阳性校正比最先进的工具(Musket,SGA,BFC,Lighter,Bcool,Karect)要少得多,同时保持了具有竞争力的真阳性校正数量。当在组装之前使用时,它可以为许多真实的数据集实现上级从头组装结果。CARE也是第一个基于多序列分析的错误校正器,它能够在单个工作站上使用GPU加速在4小时内处理人类基因组Illumina NGS数据集。
Motivation: Error correction is a fundamental pre-processing step in many Next-Generation Sequencing (NGS) pipelines, in particular for de novo genome assembly. However, existing error correction methods either suffer from high false-positive rates since they break reads into independent k-mers or do not scale efficiently to large amounts of sequencing reads and complex genomes.Results: We present CARE-an alignment-based scalable error correction algorithm for Illumina data using the concept of minhashing. Minhashing allows for efficient similarity search within large sequencing read collections which enables fast computation of high-quality multiple alignments. Sequencing errors are corrected by detailed inspection of the corresponding alignments. Our performance evaluation shows that CARE generates significantly fewer false-positive corrections than state-of-the-art tools (Musket, SGA, BFC, Lighter, Bcool, Karect) while maintaining a competitive number of true positives. When used prior to assembly it can achieve superior de novo assembly results for a number of real datasets. CARE is also the first multiple sequence alignment-based error corrector that is able to process a human genome Illumina NGS dataset in only 4 h on a single workstation using GPU acceleration.