MUMmer4: A fast and versatile genome alignment system.

MUMmer4: A fast and versatile genome alignment system.
复制标题

DOI:
10.1371/journal.pcbi.1005944
复制
发表时间:
2018-01
影响因子:
4.3
通讯作者:
Zimin A
Zimin A
中科院分区:
生物学2区
文献类型:
--
作者:
Marçais G;Delcher AL;Phillippy AM;Coston R;Salzberg SL;Zimin A

文献摘要

参考文献

被引文献

相似文献

MUMmer系统和其中包括的基因组序列比对器nucmer是基因组学中最广泛使用的比对包之一。自2004年MUMmer第3版的最后一个主要版本以来,它已被应用于许多类型的问题,包括比对全基因组序列,比对读段与参考基因组,以及比较相同基因组的不同组装。尽管MUMMer 3具有广泛的实用性,但它存在局限性,难以用于大型基因组和当今常见的非常大的序列数据集。在本文中,我们描述了MUMMer 4,一个大大改进的版本MUMMer,解决基因组大小的限制,通过改变32位的后缀树数据结构的核心MUMMer到一个48位的后缀数组,并提供了改进的速度,通过并行处理输入查询序列。由于输入长度的理论限制为141 Tbp,MUMMer 4现在可以处理任何生物学上实际长度的输入序列。我们表明,作为这些增强的结果,MUMMer 4中的nucmer程序很容易处理大型基因组的比对;我们用人类和黑猩猩基因组的比对来说明这一点,这使我们能够计算出这两个物种在96%的长度上98%相同。通过本文所述的增强,MUMmer 4也可以用于有效地将读数与参考基因组进行比对,尽管它不如专用读数比对器灵敏和准确。MUMmer 4中的nucmer aligner现在可以从Perl,Python和Ruby等脚本语言调用。这些改进使MUMer 4成为最通用的基因组比对软件包之一。
The MUMmer system and the genome sequence aligner nucmer included within it are among the most widely used alignment packages in genomics. Since the last major release of MUMmer version 3 in 2004, it has been applied to many types of problems including aligning whole genome sequences, aligning reads to a reference genome, and comparing different assemblies of the same genome. Despite its broad utility, MUMmer3 has limitations that can make it difficult to use for large genomes and for the very large sequence data sets that are common today. In this paper we describe MUMmer4, a substantially improved version of MUMmer that addresses genome size constraints by changing the 32-bit suffix tree data structure at the core of MUMmer to a 48-bit suffix array, and that offers improved speed through parallel processing of input query sequences. With a theoretical limit on the input size of 141Tbp, MUMmer4 can now work with input sequences of any biologically realistic length. We show that as a result of these enhancements, the nucmer program in MUMmer4 is easily able to handle alignments of large genomes; we illustrate this with an alignment of the human and chimpanzee genomes, which allows us to compute that the two species are 98% identical across 96% of their length. With the enhancements described here, MUMmer4 can also be used to efficiently align reads to reference genomes, although it is less sensitive and accurate than the dedicated read aligners. The nucmer aligner in MUMmer4 can now be called from scripting languages such as Perl, Python and Ruby. These improvements make MUMer4 one the most versatile genome alignment packages available.
DOI: 10.1038/nmeth.1923
发表时间: 2012-03-04
期刊: NATURE METHODS
影响因子: 48
作者:
Langmead, Ben;Salzberg, Steven L.
通讯作者: Salzberg, Steven L.
DOI: 10.1101/gr.213611.116
发表时间: 2017-05-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Schneider, Valerie A.;Graves-Lindsay, Tina;Church, Deanna M.
通讯作者: Church, Deanna M.
DOI: 10.1016/j.tcs.2007.07.017
发表时间: 2007-11-22
影响因子: 1.1
作者:
Larsson, N. Jesper;Sadakane, Kunihiko
通讯作者: Sadakane, Kunihiko
DOI: 10.1186/gb-2009-10-3-r25
发表时间: 2009
期刊: Genome biology
影响因子: 12.3
作者:
Langmead B;Trapnell C;Pop M;Salzberg SL
通讯作者: Salzberg SL
DOI: 10.1093/bioinformatics/btp324
发表时间: 2009-07-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Li H;Durbin R
通讯作者: Durbin R