HALC: High throughput algorithm for long read error correction.

HALC: High throughput algorithm for long read error correction.
复制标题

HALC:用于长读纠错的高吞吐量算法

DOI:
10.1186/s12859-017-1610-3
复制
发表时间:
2017-04-05
期刊:
影响因子:
3
通讯作者:
Lan L
Lan L
中科院分区:
生物学4区
文献类型:
--
作者:
Bao E;Lan L

文献摘要

被引文献

相似文献

第三代PacBio SMRT长读段可以有效解决第二代测序技术的读段长度问题,但包含约15%的测序错误。已经设计了几种纠错算法来有效地将错误率降低到1%,但是它们丢弃了大量未校正的碱基,从而导致低吞吐量。这种碱基损失可能会限制下游组装的完整性和分析的准确性。在这里,我们介绍HALC,一个高吞吐量的算法,用于长读错误校正。HALC以相对低的同一性要求将长读段与来自相同物种的短读段叠连群进行比对,使得长读段区域可以与至少一个叠连群区域进行比对,包括与其足够相似的叠连群中的其真实基因组区域的重复(基于相似重复的比对方法)。然后构建重叠群图,并且对于每个长读段,参考其他长读段的比对以找到最准确的比对,并使用比对的重叠群区域对其进行校正(基于长读段支持的验证方法)。即使重叠群中没有真实基因组区域的一些长读段区域用其重复序列校正,这种方法也使得可以进一步细化这些具有初始不足的短读段的长读段区域并校正其间的未校正区域。在我们对E. coli、A. thaliana和Maylandia zebra数据集,HALC能够获得比现有算法高6.7-41.1%的吞吐量,同时保持相当的准确性。因此,HALC校正的长读段可以产生比现有算法长11.4-60.7%的组装的重叠群。HALC软件可以从这个网站免费下载:https://github.com/lanl001/halc。本文的在线版本(doi:10.1186/s12859-017-1610-3)包含补充材料,可供授权用户使用。
The third generation PacBio SMRT long reads can effectively address the read length issue of the second generation sequencing technology, but contain approximately 15% sequencing errors. Several error correction algorithms have been designed to efficiently reduce the error rate to 1%, but they discard large amounts of uncorrected bases and thus lead to low throughput. This loss of bases could limit the completeness of downstream assemblies and the accuracy of analysis. Here, we introduce HALC, a high throughput algorithm for long read error correction. HALC aligns the long reads to short read contigs from the same species with a relatively low identity requirement so that a long read region can be aligned to at least one contig region, including its true genome region’s repeats in the contigs sufficiently similar to it (similar repeat based alignment approach). It then constructs a contig graph and, for each long read, references the other long reads’ alignments to find the most accurate alignment and correct it with the aligned contig regions (long read support based validation approach). Even though some long read regions without the true genome regions in the contigs are corrected with their repeats, this approach makes it possible to further refine these long read regions with the initial insufficient short reads and correct the uncorrected regions in between. In our performance tests on E. coli, A. thaliana and Maylandia zebra data sets, HALC was able to obtain 6.7-41.1% higher throughput than the existing algorithms while maintaining comparable accuracy. The HALC corrected long reads can thus result in 11.4-60.7% longer assembled contigs than the existing algorithms. The HALC software can be downloaded for free from this site: https://github.com/lanl001/halc. The online version of this article (doi:10.1186/s12859-017-1610-3) contains supplementary material, which is available to authorized users.