DBG2OLC: Efficient Assembly of Large Genomes Using Long Erroneous Reads of the Third Generation Sequencing Technologies.

DBG2OLC: Efficient Assembly of Large Genomes Using Long Erroneous Reads of the Third Generation Sequencing Technologies.
复制标题

DBG2OLC:使用第三代测序技术的长错误读取高效组装大基因组

DOI:
10.1038/srep31900
复制
发表时间:
2016-08-30
期刊:
影响因子:
4.6
通讯作者:
Ma ZS
Ma ZS
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Ye C;Hill CM;Wu S;Ruan J;Ma ZS

文献摘要

参考文献

被引文献

相似文献

从下一代测序(NGS)到第三代测序(3GS)的高度预期的过渡一直是困难的,主要是由于高错误率和过高的测序成本。高错误率使得大基因组的长错误读段的组装具有挑战性,因为现有的软件解决方案通常被错误校正任务淹没。在这里,我们报告了一种混合组装方法,同时利用NGS和3GS数据来解决这两个问题。我们从三个一般和基本的设计原则中获得优势:(i)长读段的紧凑表示导致有效的比对。(ii)基本错误可以跳过;结构性错误需要检测和纠正。(iii)结构正确的3GS读数被组装和抛光。在我们的实现中,预组装的NGS重叠群被用来获得长读段的紧凑表示,激励从ade Bruijngraph到重叠图的算法转换,这两个主要的组装范例。此外,由于NGS和3GS数据可以相互补偿,我们的混合组装方法降低了它们的测序要求。实验表明,我们的软件能够比现有的方法更快地组装多个数量级的基因组,而不会消耗大量的内存,同时节省约一半的测序成本。
The highly anticipated transition from next generation sequencing (NGS) to third generation sequencing (3GS) has been difficult primarily due to high error rates and excessive sequencing cost. The high error rates make the assembly of long erroneous reads of large genomes challenging because existing software solutions are often overwhelmed by error correction tasks. Here we report a hybrid assembly approach that simultaneously utilizes NGS and 3GS data to address both issues. We gain advantages from three general and basic design principles: (i) Compact representation of the long reads leads to efficient alignments. (ii) Base-level errors can be skipped; structural errors need to be detected and corrected. (iii) Structurally correct 3GS reads are assembled and polished. In our implementation, preassembled NGS contigs are used to derive the compact representation of the long reads, motivating an algorithmic conversion from ade Bruijngraph to an overlap graph, the two major assembly paradigms. Moreover, since NGS and 3GS data can compensate for each other, our hybrid assembly approach reduces both of their sequencing requirements. Experiments show that our software is able to assemble mammalian-sized genomes orders of magnitude more quickly than existing methods without consuming a lot of memory, while saving about half of the sequencing cost.
DOI: 10.1101/gr.141515.112
发表时间: 2012-11
期刊: Genome research
影响因子: 7
作者:
Ribeiro FJ;Przybylski D;Yin S;Sharpe T;Gnerre S;Abouelleil A;Berlin AM;Montmayeur A;Shea TP;Walker BJ;Young SK;Russ C;Nusbaum C;MacCallum I;Jaffe DB
通讯作者: Jaffe DB
DOI: 10.1186/1471-2105-13-s6-s1
发表时间: 2012-04-19
期刊: BMC bioinformatics
影响因子: 3
作者:
Ye C;Ma ZS;Cannon CH;Pop M;Yu DW
通讯作者: Yu DW
DOI: 10.1093/bioinformatics/btu538
发表时间: 2014-12-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Salmela, Leena;Rivals, Eric
通讯作者: Rivals, Eric
DOI: 10.1101/gr.079053.108
发表时间: 2009-02-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Chaisson, Mark J.;Brinza, Dumitru;Pevzner, Pavel A.
通讯作者: Pevzner, Pavel A.
DOI: 10.1093/bioinformatics/bti1114
发表时间: 2005-09-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Myers, EW
通讯作者: Myers, EW