QAlign: aligning nanopore reads accurately using current-level modeling.

QAlign: aligning nanopore reads accurately using current-level modeling.
复制标题

QAlign:使用电流水平建模准确对齐纳米孔读数。

DOI:
10.1093/bioinformatics/btaa875
复制
发表时间:
2021
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Diggavi,Suhas
Diggavi,Suhas
中科院分区:
--
文献类型:
--
作者:
Joshi,Dhaivat;Mao,Shunfu;Kannan,Sreeram;Diggavi,Suhas

文献摘要

相似文献

动机DNA/RNA 序列读数相互之间或与参考基因组/转录组的高效且准确的比对是基因组分析中的一个重要问题。纳米孔测序已成为一种主要的测序技术,并且许多长读长对齐器被设计用于对齐纳米孔读长。然而,高错误率使得准确有效的对准变得困难。正确利用测序过程中固有的噪声和错误特征可以在构建稳健的比对器中发挥至关重要的作用。在本文中,我们设计了 QAlign,这是一种预处理器,可与任何长读对齐器一起使用,用于将长读长与基因组/转录组或其他长读长对齐。 QAlign 的关键思想是将核苷酸读数转换为离散电流水平,在通过序列比对器运行之前捕获纳米孔测序仪的错误模式。结果我们表明,当与基因组比对时,QAlign 能够提高从周围到纳米孔读数的比对率。我们还表明,QAlign 通过在三个真实数据集中进行读取到读取的对齐来提高平均重叠质量。在两个真实数据集中,读取到转录组比对率从 % 提高到 % 到 % 。可用性和实现https://github.com/joshidhaivat/QAlign.git。补充信息补充数据可在 Bioinformaticsonline 上获得。
MotivationEfficient and accurate alignment of DNA/RNA sequence reads to each other or to a reference genome/transcriptome is an important problem in genomic analysis. Nanopore sequencing has emerged as a major sequencing technology and many long-read aligners have been designed for aligning nanopore reads. However, the high error rate makes accurate and efficient alignment difficult. Utilizing the noise and error characteristics inherent in the sequencing process properly can play a vital role in constructing a robust aligner. In this article, we design QAlign, a pre-processor that can be used with any long-read aligner for aligning long reads to a genome/transcriptome or to other long reads. The key idea in QAlign is to convert the nucleotide reads into discretized current levels that capture the error modes of the nanopore sequencer before running it through a sequence aligner.ResultsWe show that QAlign is able to improve alignment rates from aroundup towith nanopore reads when aligning to the genome. We also show that QAlign improves the average overlap quality byandin three real datasets for read-to-read alignment. Read-to-transcriptome alignment rates are improved from% toand% toin two real datasets.Availability and implementationhttps://github.com/joshidhaivat/QAlign.git.Supplementary informationSupplementary data are available atBioinformaticsonline.