NanoReviser: An Error-Correction Tool for Nanopore Sequencing Based on a Deep Learning Algorithm

NanoReviser: An Error-Correction Tool for Nanopore Sequencing Based on a Deep Learning Algorithm
复制标题

NanoReviser:基于深度学习算法的纳米孔测序纠错工具。

DOI:
10.3389/fgene.2020.00900
复制
发表时间:
2020-08-12
影响因子:
3.7
通讯作者:
Zhu, Huaiqiu
Zhu, Huaiqiu
中科院分区:
生物学3区
文献类型:
--
作者:
Wang, Luotong;Qu, Li;Zhu, Huaiqiu

文献摘要

被引文献

相似文献

纳米孔测序被认为是最有前途的第三代测序(TGS)技术之一。自2014年以来,Oxford Nanopore Technologies(ONT)开发了一系列基于纳米孔测序的设备,以产生非常长的读数,对基因组学产生了预期的影响。然而,由于难以从复杂的电信号中识别DNA碱基,纳米孔测序读数容易受到相当高的错误率的影响。尽管在过去的几年中已经开发了几种碱基识别工具用于纳米孔测序,但在应用碱基识别程序后校正序列仍然具有挑战性。在这项研究中,我们开发了一个开源的DNA碱基判定修订器NanoReviser,它基于一种深度学习算法来纠正默认提供的当前碱基判定器引入的碱基判定错误。在我们的模块中,我们根据默认碱基检测器提供的碱基检测序列重新分割原始电信号。通过使用卷积神经网络(CNN)和双向长短期记忆(Bi-LSTM)网络,我们利用了来自原始电信号和来自碱基识别器的碱基识别序列的信息。我们的研究结果表明,NanoReviser作为碱基判定后的修订者,显著提高了碱基判定质量。在对来自publicE.对于大肠杆菌和人类NA 12878数据集,NanoReviser将两者的测序错误率降低了5%以上。colisetaset和human dataset。NanoReviser的性能被发现优于所有当前碱基判定工具。此外,我们还分析了E. colidataset并添加了甲基化信息来训练我们的模块。通过甲基化注释,NanoReviser将E. colidataset,特别是减少了超过10%的错误率的区域序列丰富的甲基化碱基。据我们所知,NanoReviser是碱基识别后的第一个后处理工具,可以准确校正纳米孔序列,而无需耗时的构建共有序列的过程。NanoReviser软件包可从免费获得。https://github.com/pkubioinformatics/NanoReviser.
Nanopore sequencing is regarded as one of the most promising third-generation sequencing (TGS) technologies. Since 2014, Oxford Nanopore Technologies (ONT) has developed a series of devices based on nanopore sequencing to produce very long reads, with an expected impact on genomics. However, the nanopore sequencing reads are susceptible to a fairly high error rate owing to the difficulty in identifying the DNA bases from the complex electrical signals. Although several basecalling tools have been developed for nanopore sequencing over the past years, it is still challenging to correct the sequences after applying the basecalling procedure. In this study, we developed an open-source DNA basecalling reviser, NanoReviser, based on a deep learning algorithm to correct the basecalling errors introduced by current basecallers provided by default. In our module, we re-segmented the raw electrical signals based on the basecalled sequences provided by the default basecallers. By employing convolution neural networks (CNNs) and bidirectional long short-term memory (Bi-LSTM) networks, we took advantage of the information from the raw electrical signals and the basecalled sequences from the basecallers. Our results showed NanoReviser, as a post-basecalling reviser, significantly improving the basecalling quality. After being trained on standard ONT sequencing reads from publicE. coliand human NA12878 datasets, NanoReviser reduced the sequencing error rate by over 5% for both theE. colidataset and the human dataset. The performance of NanoReviser was found to be better than those of all current basecalling tools. Furthermore, we analyzed the modified bases of theE. colidataset and added the methylation information to train our module. With the methylation annotation, NanoReviser reduced the error rate by 7% for theE. colidataset and specifically reduced the error rate by over 10% for the regions of the sequence rich in methylated bases. To the best of our knowledge, NanoReviser is the first post-processing tool after basecalling to accurately correct the nanopore sequences without the time-consuming procedure of building the consensus sequence. The NanoReviser package is freely available at. https://github.com/pkubioinformatics/NanoReviser.