BlockPolish: accurate polishing of long-read assembly via block divide-and-conquer

BlockPolish: accurate polishing of long-read assembly via block divide-and-conquer
复制标题

DOI:
10.1093/bib/bbab405
复制
发表时间:
2021-10
影响因子:
9.5
通讯作者:
Neng Huang;Fan Nie;Peng Ni;Xin Gao;F. Luo;Jianxin Wang
Neng Huang;Fan Nie;Peng Ni;Xin Gao;F. Luo;Jianxin Wang
中科院分区:
生物学2区
文献类型:
--
作者:
Neng Huang;Fan Nie;Peng Ni;Xin Gao;F. Luo;Jianxin Wang

文献摘要

相似文献

长读测序技术使从头基因组组装取得重大进展。然而,由于原始读取的高错误率和广泛的错误分布,导致汇编中存在大量错误。抛光是一个程序,以解决错误的草案组装和提高基因组分析的可靠性。然而,现有的方法对装配体的所有区域都一视同仁,而这些区域的误差分布存在根本的差异。如何在基因组组装中达到很高的精度仍然是一个具有挑战性的问题。针对装配体不同区域的不均匀误差,提出了一种新的抛光工作流程——BlockPolish。在该方法中,我们根据排列核苷酸碱基的统计数据将contigs划分为低复杂度和高复杂度的块。采用多序列比对技术对复杂块中的原始读取进行重新比对,优化比对结果。针对简单块和复杂块错误率分布的不同,提出了两种多任务双向长短期记忆(LSTM)网络来预测一致性序列。在Wtdbg2和Flye利用纳米孔数据组装的NA12878全基因组片段中,BlockPolish的抛光精度高于Racon、Medaka和MarginPolish & HELEN等其他最先进的工具。在所有程序集中,错误主要是索引,BlockPolish在纠正错误方面表现良好。除了纳米孔组件,我们进一步证明BlockPolish还可以减少PacBio组件中的误差。BlockPolish的源代码可以在Github上免费获得(https://github.com/huangnengCSU/BlockPolish)。
Long-read sequencing technology enables significant progress in de novo genome assembly. However, the high error rate and the wide error distribution of raw reads result in a large number of errors in the assembly. Polishing is a procedure to fix errors in the draft assembly and improve the reliability of genomic analysis. However, existing methods treat all the regions of the assembly equally while there are fundamental differences between the error distributions of these regions. How to achieve very high accuracy in genome assembly is still a challenging problem. Motivated by the uneven errors in different regions of the assembly, we propose a novel polishing workflow named BlockPolish. In this method, we divide contigs into blocks with low complexity and high complexity according to statistics of aligned nucleotide bases. Multiple sequence alignment is applied to realign raw reads in complex blocks and optimize the alignment result. Due to the different distributions of error rates in trivial and complex blocks, two multitask bidirectional Long short-term memory (LSTM) networks are proposed to predict the consensus sequences. In the whole-genome assemblies of NA12878 assembled by Wtdbg2 and Flye using Nanopore data, BlockPolish has a higher polishing accuracy than other state-of-the-arts including Racon, Medaka and MarginPolish & HELEN. In all assemblies, errors are predominantly indels and BlockPolish has a good performance in correcting them. In addition to the Nanopore assemblies, we further demonstrate that BlockPolish can also reduce the errors in the PacBio assemblies. The source code of BlockPolish is freely available on Github (https://github.com/huangnengCSU/BlockPolish).