Safe and Complete Contig Assembly Through Omnitigs

Safe and Complete Contig Assembly Through Omnitigs
复制标题

DOI:
10.1089/cmb.2016.0141
复制
发表时间:
2017-06-01
影响因子:
1.7
通讯作者:
Medvedev, Paul
Medvedev, Paul
中科院分区:
生物学4区
文献类型:
--
作者:
Tomescu, Alexandru I.;Medvedev, Paul

文献摘要

被引文献

相似文献

从一组reads重构基因组时,cong组装是大多数组装者要解决的第一个问题。它的输出由一组串组成,这些串被承诺出现在任何可能产生读取的基因组中。自20年前引入contigs以来,组装者一直试图获得越来越长的contigs,但以下问题仍然存在:给定一个基因组图G(例如,de Bruijn或字符串图),可以安全地从G报告为contigs的所有字符串是什么?在这篇文章中,我们使用基因组是一个圆形覆盖行走的模型来回答这个问题。我们还给出了一个多项式时间算法来找到这样的字符串,我们称之为omnitigs。我们的实验表明,omnitigs比流行的单元平均长66%-82%,并且29%的dbSNP位置在omnitigs中比在unigs中有更多的邻居。
Contig assembly is the first stage that most assemblers solve when reconstructing a genome from a set of reads. Its output consists of contigsa set of strings that are promised to appear in any genome that could have generated the reads. From the introduction of contigs 20 years ago, assemblers have tried to obtain longer and longer contigs, but the following question remains: given a genome graph G (e.g., a de Bruijn, or a string graph), what are all the strings that can be safely reported from G as contigs? In this article, we answer this question using a model in which the genome is a circular covering walk. We also give a polynomial-time algorithm to find such strings, which we call omnitigs. Our experiments show that omnitigs are 66%-82% longer on average than the popular unitigs, and 29% of dbSNP locations have more neighbors in omnitigs than in unitigs.