SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler.

SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler.
复制标题

DOI:
10.1186/2047-217x-1-18
复制
发表时间:
2012-12-27
期刊:
影响因子:
9.2
通讯作者:
Wang J
Wang J
中科院分区:
生物学2区
文献类型:
--
作者:
Luo R;Liu B;Xie Y;Li Z;Huang W;Yuan J;He G;Chen Y;Pan Q;Liu Y;Tang J;Wu G;Zhang H;Shi Y;Liu Y;Yu C;Wang B;Lu Y;Han C;Cheung DW;Yiu SM;Peng S;Xiaoqian Z;Liu G;Liao X;Li Y;Yang H;Wang J;Lam TW;Wang J

文献摘要

被引文献

相似文献

使用下一代测序(NGS)短读段的从头基因组组装的数量迅速增加;然而,为了使其高效和准确,仍需克服几个重大挑战。SOAPDenovo已经成功地应用于许多已发表的基因组的组装,但它仍然需要改进的连续性,准确性和覆盖率,特别是在重复区域。为了克服这些挑战,我们开发了它的继任者SOAPDenovo 2,它具有新算法设计的优势,可以减少图形构建中的内存消耗,解决重叠群组装中的更多重复区域,增加支架构建中的覆盖范围和长度,改善缺口闭合,并优化大型基因组。使用Assemblathon 1和GAGE数据集进行的基准测试表明,SOAPDenovo 2大大超过了其前身SOAPDenovo,并且在装配长度和精度方面与其他装配器具有竞争力。我们还提供了一个更新的汇编版本的2008年亚洲(YH)基因组使用SOAPDenovo 2。在这里,YH基因组的重叠群和支架N50分别为~20.9 kbp和~22 Mbp,这是第一个公开版本的3倍和50倍长。基因组覆盖率从81.16%提高到93.91%,在最大内存消耗点时,内存消耗降低了约2/3。
There is a rapidly increasing amount of de novo genome assembly using next-generation sequencing (NGS) short reads; however, several big challenges remain to be overcome in order for this to be efficient and accurate. SOAPdenovo has been successfully applied to assemble many published genomes, but it still needs improvement in continuity, accuracy and coverage, especially in repeat regions. To overcome these challenges, we have developed its successor, SOAPdenovo2, which has the advantage of a new algorithm design that reduces memory consumption in graph construction, resolves more repeat regions in contig assembly, increases coverage and length in scaffold construction, improves gap closing, and optimizes for large genome. Benchmark using the Assemblathon1 and GAGE datasets showed that SOAPdenovo2 greatly surpasses its predecessor SOAPdenovo and is competitive to other assemblers on both assembly length and accuracy. We also provide an updated assembly version of the 2008 Asian (YH) genome using SOAPdenovo2. Here, the contig and scaffold N50 of the YH genome were ~20.9 kbp and ~22 Mbp, respectively, which is 3-fold and 50-fold longer than the first published version. The genome coverage increased from 81.16% to 93.91%, and memory consumption was ~2/3 lower during the point of largest memory consumption.