HQAlign: aligning nanopore reads for SV detection using current-level modeling

HQAlign: aligning nanopore reads for SV detection using current-level modeling
复制标题

HQAalign:使用当前水平建模对齐纳米孔读数以进行 SV 检测

DOI:
10.1093/bioinformatics/btad580
复制
发表时间:
2023
期刊:
影响因子:
5.8
通讯作者:
Kannan, Sreeram
Kannan, Sreeram
中科院分区:
生物学3区
文献类型:
--
作者:
Joshi, Dhaivat;Diggavi, Suhas;Chaisson, Mark J;Kannan, Sreeram

文献摘要

相似文献

从样品DNA读数与参考基因组的比对中检测结构变异(SV)是理解人类疾病的重要问题。可以跨越重复区域的长读段,沿着这些长读段的精确比对,在鉴定新SV中起重要作用。长读段测序仪,如纳米孔测序,可以通过提供非常长的读段来解决这个问题,但错误率很高,这使得精确比对具有挑战性。由纳米孔测序引起的许多误差由于测序过程的物理性质而具有偏差,并且这些误差特性的适当利用可以在设计用于SV检测问题的稳健比对器中发挥重要作用。在这篇文章中,我们设计和评估HQAlign,一种使用纳米孔测序读数进行SV检测的比对器。HQAlign的关键思想包括(i)使用称为纳米孔读取的碱基沿着纳米孔物理学来改善SV的比对,(ii)将SV特异性变化并入比对流水线,以及(iii)将这些调整到现有的最先进的长读取比对流水线minimap 2中(v2.24),用于有效的比对。结果我们表明,HQAlign在不同的数据集上捕获了大约4%-6%的互补SV,其被Minimap 2比对错过,同时对于真实的纳米孔读取数据具有与Minimap 2相当的独立性能。对于HQAlign和minimap 2之间的常见SV调用,HQAlign将不同数据集上SV的开始和结束断点准确度提高了约10%-50%。此外,HQAlign将纳米孔读数比对到最近的端粒到端粒CHM 13组装的比对率从minimap 2的85.64%提高到89.35%,并且将纳米孔读数比对到GRCh 37人类基因组的比对率从83.48%提高到86.65%。github.com/joshidhaivat/HQAlign.git
MotivationDetection of structural variants (SVs) from the alignment of sample DNA reads to the reference genome is an important problem in understanding human diseases. Long reads that can span repeat regions, along with an accurate alignment of these long reads play an important role in identifying novel SVs. Long-read sequencers, such as nanopore sequencing, can address this problem by providing very long reads but with high error rates, making accurate alignment challenging. Many errors induced by nanopore sequencing have a bias because of the physics of the sequencing process and proper utilization of these error characteristics can play an important role in designing a robust aligner for SV detection problems. In this article, we design and evaluate HQAlign, an aligner for SV detection using nanopore sequenced reads. The key ideas of HQAlign include (i) using base-called nanopore reads along with the nanopore physics to improve alignments for SVs, (ii) incorporating SV-specific changes to the alignment pipeline, and (iii) adapting these into existing state-of-the-art long-read aligner pipeline, minimap2 (v2.24), for efficient alignments.ResultsWe show that HQAlign captures about 4%–6% complementary SVs across different datasets, which are missed by minimap2 alignments while having a standalone performance at par with minimap2 for real nanopore reads data. For the common SV calls between HQAlign and minimap2, HQAlign improves the start and the end breakpoint accuracy by about 10%–50% for SVs across different datasets. Moreover, HQAlign improves the alignment rate to 89.35% from minimap2 85.64% for nanopore reads alignment to recent telomere-to-telomere CHM13 assembly, and it improves to 86.65% from 83.48% for nanopore reads alignment to GRCh37 human genome.Availability and implementationhttps://github.com/joshidhaivat/HQAlign.git.