LINKS: Scalable, alignment-free scaffolding of draft genomes with long reads.

LINKS: Scalable, alignment-free scaffolding of draft genomes with long reads.
复制标题

DOI:
10.1186/s13742-015-0076-3
复制
发表时间:
2015
期刊:
影响因子:
9.2
通讯作者:
Birol I
Birol I
中科院分区:
生物学2区
文献类型:
--
作者:
Warren RL;Yang C;Vandervalk BP;Behsaz B;Lagman A;Jones SJ;Birol I

文献摘要

被引文献

相似文献

由于组装问题的复杂性,我们还没有完整的基因组序列。序列重复和短读段无法捕获足够的基因组信息来解决这些有问题的区域,加剧了将读段组装成最终基因组的难度。在这方面,已建立的和新兴的长读技术显示出巨大的前景,但它们当前相关的较高错误率通常需要计算碱基校正和/或额外的生物信息学预处理才能发挥价值。我们提出了 LINKS,即长间隔核苷酸 K-mer 支架算法,该方法利用纳米孔序列数据和其他包含错误的序列数据的序列特性来构建高质量的基因组组装体,无需进行读取比对或碱基校正。在这里,我们展示了如何通过使用 beta 发布的 Oxford Nanopore Technologies Ltd. 长读长将 ABySS 大肠杆菌 K-12 基因组组装的连续性提高五倍以上,以及 LINKS 如何利用酿酒酵母 W303 纳米孔读长中的长程信息来产生组装,其产生的连续性和正确性与竞争应用相当或更好。我们还展示了巨大的白云杉(Picea glauca)草图组装(PG29,20 Gbp)的重新搭建,并演示了 LINKS 如何扩展到更大的基因组。这项研究强调了纳米孔读数目前在基因组支架中的实用性,尽管它们目前存在局限性,随着纳米孔测序技术的进步,这些局限性预计会减少。我们期望 LINKS 在利用长读长连接小型和大型基因组组装草稿的高质量序列方面具有广泛的实用性。本文的在线版本 (doi:10.1186/s13742-015-0076-3) 包含补充材料,可供授权用户使用。
Owing to the complexity of the assembly problem, we do not yet have complete genome sequences. The difficulty in assembling reads into finished genomes is exacerbated by sequence repeats and the inability of short reads to capture sufficient genomic information to resolve those problematic regions. In this regard, established and emerging long read technologies show great promise, but their current associated higher error rates typically require computational base correction and/or additional bioinformatics pre-processing before they can be of value. We present LINKS, the Long Interval Nucleotide K-mer Scaffolder algorithm, a method that makes use of the sequence properties of nanopore sequence data and other error-containing sequence data, to scaffold high-quality genome assemblies, without the need for read alignment or base correction. Here, we show how the contiguity of an ABySS Escherichia coli K-12 genome assembly can be increased greater than five-fold by the use of beta-released Oxford Nanopore Technologies Ltd. long reads and how LINKS leverages long-range information in Saccharomyces cerevisiae W303 nanopore reads to yield assemblies whose resulting contiguity and correctness are on par with or better than that of competing applications. We also present the re-scaffolding of the colossal white spruce (Picea glauca) draft assembly (PG29, 20 Gbp) and demonstrate how LINKS scales to larger genomes. This study highlights the present utility of nanopore reads for genome scaffolding in spite of their current limitations, which are expected to diminish as the nanopore sequencing technology advances. We expect LINKS to have broad utility in harnessing the potential of long reads in connecting high-quality sequences of small and large genome assembly drafts. The online version of this article (doi:10.1186/s13742-015-0076-3) contains supplementary material, which is available to authorized users.