SOPRA: Scaffolding algorithm for paired reads via statistical optimization.

SOPRA: Scaffolding algorithm for paired reads via statistical optimization.
复制标题

DOI:
10.1186/1471-2105-11-345
复制
发表时间:
2010-06-24
期刊:
影响因子:
3
通讯作者:
Sengupta AM
Sengupta AM
中科院分区:
生物学4区
文献类型:
--
作者:
Dayarian A;Michael TP;Sengupta AM

文献摘要

参考文献

被引文献

相似文献

高通量测序(HTS)平台每次运行产生千兆字节的短读(<100个碱基)数据。虽然这些短的读出对于重新测序应用是足够的,但从这样的读出中重新组装中等大小的基因组仍然是一个巨大的挑战。利用配对技术可以部分克服这些限制,这种技术提供了沿基因组相隔一段已知距离的短阅读对。我们已经开发了Sopra,这是一个旨在利用配对/配对末端信息来组装短读取的工具。该算法的主要关注点是选择一个足够大的同时可满足的配对约束子集,以在输出支架的大小和质量之间实现平衡。脚手架装配被描述为与重叠连通图的顶点和边相关联的变量的优化问题。此图的顶点是由配对连接的重叠群之间绘制的边的单个重叠群。在上一代测序项目的鸟枪测序和脚手架构建的背景下,也出现了类似的图形问题。然而,考虑到HTS数据容易出错的性质以及读取时间短的根本限制,早期研究中使用的特别贪婪算法在当前情况下可能导致较差的结果质量。Sopra通过平等对待所有约束来解决优化问题来规避这个问题,解决方案本身指示有问题的约束(嵌合/重复重叠等)。将被移除。迭代求解和移除约束的过程,直到达到一致约束的核心集合。对于立体定序器数据,Sopra使用动态编程方法将颜色空间集合稳健地转换为基本空间。为了评估组件的质量,我们报告了不匹配/不匹配错误率以及各种重排错误率。将Sopra应用于来自细菌基因组的真实数据,我们能够将重叠群组装成相当长的支架(N50到200KB),在这个过程中引入的错误非常少。总体而言,这里提出的方法将允许任何类型的配对测序数据更好地组装支架。
High throughput sequencing (HTS) platforms produce gigabases of short read (<100 bp) data per run. While these short reads are adequate for resequencing applications, de novo assembly of moderate size genomes from such reads remains a significant challenge. These limitations could be partially overcome by utilizing mate pair technology, which provides pairs of short reads separated by a known distance along the genome. We have developed SOPRA, a tool designed to exploit the mate pair/paired-end information for assembly of short reads. The main focus of the algorithm is selecting a sufficiently large subset of simultaneously satisfiable mate pair constraints to achieve a balance between the size and the quality of the output scaffolds. Scaffold assembly is presented as an optimization problem for variables associated with vertices and with edges of the contig connectivity graph. Vertices of this graph are individual contigs with edges drawn between contigs connected by mate pairs. Similar graph problems have been invoked in the context of shotgun sequencing and scaffold building for previous generation of sequencing projects. However, given the error-prone nature of HTS data and the fundamental limitations from the shortness of the reads, the ad hoc greedy algorithms used in the earlier studies are likely to lead to poor quality results in the current context. SOPRA circumvents this problem by treating all the constraints on equal footing for solving the optimization problem, the solution itself indicating the problematic constraints (chimeric/repetitive contigs, etc.) to be removed. The process of solving and removing of constraints is iterated till one reaches a core set of consistent constraints. For SOLiD sequencer data, SOPRA uses a dynamic programming approach to robustly translate the color-space assembly to base-space. For assessing the quality of an assembly, we report the no-match/mismatch error rate as well as the rates of various rearrangement errors. Applying SOPRA to real data from bacterial genomes, we were able to assemble contigs into scaffolds of significant length (N50 up to 200 Kb) with very few errors introduced in the process. In general, the methodology presented here will allow better scaffold assemblies of any type of mate pair sequencing data.
来自非常短的读物的新型细菌基因组的基因促进组装。
DOI: 10.1371/journal.pcbi.1000186
发表时间: 2008-09-26
影响因子: 4.3
作者:
Salzberg, Steven L.;Sommer, Daniel D.;Puiu, Daniela;Lee, Vincent T.
通讯作者: Lee, Vincent T.
DOI: 10.1101/gr.7088808
发表时间: 2008-02-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Chaisson, Mark J.;Pevzner, Pavel A.
通讯作者: Pevzner, Pavel A.
DOI: 10.1111/j.1574-6968.2008.01441.x
发表时间: 2009-02-01
影响因子: 2.1
作者:
Farrer, Rhys A.;Kemen, Eric;Studholme, David J.
通讯作者: Studholme, David J.
DOI: 10.1088/0305-4470/15/10/028
发表时间: 1982-01-01
期刊: JOURNAL OF PHYSICS A-MATHEMATICAL AND GENERAL
影响因子: --
作者:
BARAHONA, F
通讯作者: BARAHONA, F
DOI: 10.1093/bioinformatics/btm451
发表时间: 2007-11-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Jeck, William R.;Reinhardt, Josephine A.;Jones, Corbin D.
通讯作者: Jones, Corbin D.