A hybrid and scalable error correction algorithm for indel and substitution errors of long reads

A hybrid and scalable error correction algorithm for indel and substitution errors of long reads
复制标题

DOI:
10.1186/s12864-019-6286-9
复制
发表时间:
2019-12-20
期刊:
影响因子:
4.4
通讯作者:
Park, Seung-Jong
Park, Seung-Jong
中科院分区:
生物学2区
文献类型:
--
作者:
Das, Arghya Kusum;Goswami, Sayan;Park, Seung-Jong

文献摘要

被引文献

相似文献

背景资料:长读段测序已经显示出通过提供更完整的组装来克服第二代测序的短长度限制的前景。然而,长测序读段的计算受到其较高错误率的挑战(例如,13%与1%)和更高的成本(0.3美元与0.03美元每Mbp)相比,短reads.Methods:在本文中,我们提出了一种新的混合纠错工具,称为ParLECH(并行长读纠错使用混合方法)。ParLECH的纠错算法本质上是分布式的,并且有效地利用高通量Illumina短读段序列的k-mer覆盖信息来纠正PacBio长读段序列。ParLECH首先从短读段构建de Bruijn图,然后在基于短读段的de Bruijn图中将长读段的indel错误区域替换为它们相应的最宽路径(或最大最小覆盖路径)。ParLECH然后利用短读段的k聚体覆盖信息将每个长读段分成低覆盖区域和高覆盖区域的序列,然后通过多数表决来纠正每个替换的错误bases.Results:ParLECH优于最新的最先进的混合错误校正方法对真实的PacBio数据集。我们的实验评估结果表明,ParLECH可以纠正大规模的真实世界的数据集在一个准确的和可扩展的方式。ParLECH可以使用128个计算节点在不到29小时内纠正人类基因组PacBio长读段(312 GB)与Illumina短读段(452 GB)的indel错误。ParLECH可以比对大肠杆菌92%以上的碱基。coliPacBio数据集与参考基因组的比对,证明了其准确性。结论:ParLECH可以使用数百个计算节点扩展到超过TB的测序数据。所提出的混合纠错方法是新颖的,并且纠正了原始长读段中存在的或由短读段新引入的插入缺失和置换错误。
Background: Long-read sequencing has shown the promises to overcome the short length limitations of second-generation sequencing by providing more complete assembly. However, the computation of the long sequencing reads is challenged by their higher error rates (e.g., 13% vs. 1%) and higher cost ($0.3 vs. $0.03 per Mbp) compared to the short reads.Methods: In this paper, we present a new hybrid error correction tool, called ParLECH (Parallel Long-read Error Correction using Hybrid methodology). The error correction algorithm of ParLECH is distributed in nature and efficiently utilizes the k-mer coverage information of high throughput Illumina short-read sequences to rectify the PacBio long-read sequences. ParLECH first constructs a de Bruijn graph from the short reads, and then replaces the indel error regions of the long reads with their corresponding widest path (or maximum min-coverage path) in the short read-based de Bruijn graph. ParLECH then utilizes the k-mer coverage information of the short reads to divide each long read into a sequence of low and high coverage regions, followed by a majority voting to rectify each substituted error base.Results: ParLECH outperforms latest state-of-the-art hybrid error correction methods on real PacBio datasets. Our experimental evaluation results demonstrate that ParLECH can correct large-scale real-world datasets in an accurate and scalable manner. ParLECH can correct the indel errors of human genome PacBio long reads (312 GB) with Illumina short reads (452 GB) in less than 29 h using 128 compute nodes. ParLECH can align more than 92% bases of an E. coli PacBio dataset with the reference genome, proving its accuracy.Conclusion: ParLECH can scale to over terabytes of sequencing data using hundreds of computing nodes. The proposed hybrid error correction methodology is novel and rectifies both indel and substitution errors present in the original long reads or newly introduced by the short reads.