lra: A long read aligner for sequences and contigs.

lra: A long read aligner for sequences and contigs.
复制标题

LRA:用于序列和重叠群的长读码校准器。

DOI:
10.1371/journal.pcbi.1009078
复制
发表时间:
2021-06
影响因子:
4.3
通讯作者:
Chaisson MJP
Chaisson MJP
中科院分区:
生物学2区
文献类型:
--
作者:
Ren J;Chaisson MJP

文献摘要

参考文献

被引文献

相似文献

通过对齐单分子测序(SMS)读段或来自SMS组装的contigs来检测变异在计算上具有挑战性。有效地对齐SMS读取的一种方法是稀疏动态规划(SDP),其中在序列和基因组之间找到精确匹配的最佳链。虽然直接实现SDP的代价是间隙长度的线性函数,但当间隙成本是间隙长度的凹函数时,生物变异更准确地表示出来。我们开发了一种方法lra,该方法使用具有凹成本间隙惩罚的SDP,并使用lra对PacBio和Oxford Nanopore (ONT)仪器以及从头组装的长读序列进行比对。这种比对方法提高了SV发现的敏感性和特异性,特别是对于大于1kb的变异和从ONT读取发现变异时,同时运行时间与当前方法相当(1.05-3.76倍)。当应用于从从头组装配置调用变异时,与minimap2+htsbox相比,Truvari F1得分增加了3.2%。lra可在bioconda (https://anaconda.org/bioconda/lra)和github (https://github.com/ChaissonLab/LRA)中获得。任何两个人类基因组都会在多个尺度上存在序列差异:从单核苷酸变异到大的增益、损失,或者被称为结构变异的DNA重排。长读单分子测序已被证明有助于发现结构变异,因为读段跨越了整个变异。发现结构变异的计算问题是找到读取到基因组的最佳排列,并精确地反映变异的间隙。在这里,我们展示了一种方法,lra,它使用了一个有效的实现凹成本对齐的结构变体发现使用长读取。在标准化的基准数据上,我们表明,与现有方法相比,使用lra生成的比对,对于变体检测算法和长读序列的多种组合,结构变体发现得到了改进。最后,我们证明使用lra可以使用从长读序列数据构建的从头组装准确地发现完整的结构变体谱。这意味着一个未来的比较基因组学模型,在这个模型中,只有通过比较从头组装,而不是通过比较与参考的读数,才能发现变异。
It is computationally challenging to detect variation by aligning single-molecule sequencing (SMS) reads, or contigs from SMS assemblies. One approach to efficiently align SMS reads is sparse dynamic programming (SDP), where optimal chains of exact matches are found between the sequence and the genome. While straightforward implementations of SDP penalize gaps with a cost that is a linear function of gap length, biological variation is more accurately represented when gap cost is a concave function of gap length. We have developed a method, lra, that uses SDP with a concave-cost gap penalty, and used lra to align long-read sequences from PacBio and Oxford Nanopore (ONT) instruments as well as de novo assembly contigs. This alignment approach increases sensitivity and specificity for SV discovery, particularly for variants above 1kb and when discovering variation from ONT reads, while having runtime that are comparable (1.05-3.76×) to current methods. When applied to calling variation from de novo assembly contigs, there is a 3.2% increase in Truvari F1 score compared to minimap2+htsbox. lra is available in bioconda (https://anaconda.org/bioconda/lra) and github (https://github.com/ChaissonLab/LRA). Any two human genomes will have sequence differences across multiple scales: from single-nucleotide variants to large gains, losses, or rearrangements of DNA called structural variants. Long-read single-molecule sequencing has been shown to help discover structural variation because the reads span across the entire variant. The computational problem for discovering a structural variant is to find the optimal alignment of the read to the genome with gaps that accurately reflect the variant. Here we demonstrate a method, lra, that uses an efficient implementation of concave-cost alignment for structural variant discovery using long reads. On standardized benchmark data, we show that structural variant discovery is improved for multiple combinations of variant detection algorithms and long-read sequence using alignments generated by lra compared to existing methods. Finally, we show that it is possible to use lra to accurately discover a complete spectrum of structural variants using de novo assemblies constructed from long-read sequence data. This implies a future model of comparative genomics where variants are discovered only by comparing de novo assemblies and not a comparison of reads against a reference.
DOI: 10.1038/s41592-020-01056-5
发表时间: 2021-03
期刊: Nature methods
影响因子: 48
作者:
Cheng H;Concepcion GT;Feng X;Zhang H;Li H
通讯作者: Li H
DOI: 10.1038/nature15394
发表时间: 2015-10-01
期刊: Nature
影响因子: 64.8
作者:
Sudmant PH;Rausch T;Gardner EJ;Handsaker RE;Abyzov A;Huddleston J;Zhang Y;Ye K;Jun G;Fritz MH;Konkel MK;Malhotra A;Stütz AM;Shi X;Casale FP;Chen J;Hormozdiari F;Dayama G;Chen K;Malig M;Chaisson MJP;Walter K;Meiers S;Kashin S;Garrison E;Auton A;Lam HYK;Mu XJ;Alkan C;Antaki D;Bae T;Cerveira E;Chines P;Chong Z;Clarke L;Dal E;Ding L;Emery S;Fan X;Gujral M;Kahveci F;Kidd JM;Kong Y;Lameijer EW;McCarthy S;Flicek P;Gibbs RA;Marth G;Mason CE;Menelaou A;Muzny DM;Nelson BJ;Noor A;Parrish NF;Pendleton M;Quitadamo A;Raeder B;Schadt EE;Romanovitch M;Schlattl A;Sebra R;Shabalin AA;Untergasser A;Walker JA;Wang M;Yu F;Zhang C;Zhang J;Zheng-Bradley X;Zhou W;Zichner T;Sebat J;Batzer MA;McCarroll SA;1000 Genomes Project Consortium;Mills RE;Gerstein MB;Bashir A;Stegle O;Devine SE;Lee C;Eichler EE;Korbel JO
通讯作者: Korbel JO
DOI: 10.1038/s41592-018-0001-7
发表时间: 2018-06
期刊: Nature methods
影响因子: 48
作者:
Sedlazeck FJ;Rescheneder P;Smolka M;Fang H;Nattestad M;von Haeseler A;Schatz MC
通讯作者: Schatz MC
DOI: 10.1145/146637.146650
发表时间: 1992-07-01
期刊: JOURNAL OF THE ACM
影响因子: 2.5
作者:
EPPSTEIN, D;GALIL, Z;ITALIANO, GF
通讯作者: ITALIANO, GF
DOI: 10.1371/journal.pcbi.1005944
发表时间: 2018-01
影响因子: 4.3
作者:
Marçais G;Delcher AL;Phillippy AM;Coston R;Salzberg SL;Zimin A
通讯作者: Zimin A