Assembly of repetitive regions using next-generation sequencing data

Assembly of repetitive regions using next-generation sequencing data
复制标题

DOI:
10.1016/j.bbe.2014.12.001
复制
发表时间:
2015-01-01
影响因子:
6.4
通讯作者:
Nowak, Robert M.
Nowak, Robert M.
中科院分区:
工程技术2区
文献类型:
--
作者:
Nowak, Robert M.

文献摘要

被引文献

相似文献

高读取深度可用于组装短序列重复序列。现有的基因组拼接器在长于平均读数的重复区域失败。我提出了一种新的DNA拼接算法,该算法利用相对读出频率来正确地重构重复序列。无错误输入数据的数学模型显示了作为读取覆盖率的函数的结果的精度上限。对于高覆盖,估计误差与重复序列长度成线性关系,与序列覆盖成反比。该模型描述的De Bruijn图维度越小,长重复区域的拼接就越准确,该算法需要较高的读取深度,由下一代测序仪提供,并且可以利用现有的数据。对来自几个模型基因组的无错误读取的测试,指出了正确重组的重复序列,而现有的汇编程序无法做到这一点。C++源代码、PYTHON脚本和其他数据可以在http://dnaasm.sourceforge.net.上获得(C)2014年纳莱茨生物控制论和生物医学工程研究所。作者:Elsevier Sp.ZO.O..版权所有。
High read depth can be used to assemble short sequence repeats. The existing genome assemblers fail in repetitive regions of longer than average read.I propose a new algorithm for a DNA assembly which uses the relative frequency of reads to properly reconstruct repetitive sequences. The mathematical model for error-free input data shows the upper limits of accuracy of the results as a function of read coverage. For high coverage, the estimation error depends linearly on repetitive sequence length and inversely proportional to the sequencing coverage. The model depicts, the smaller de Bruijn graph dimensions, the more accurate assembly of long repetitive regions.The algorithm requires high read depth, provided by the next-generation sequencers and could use the existing data. The tests on errorless reads, generated in silico from several model genomes, pointed the properly reconstructed repetitive sequences, where existing assemblers fail.The C++ sources, the Python scripts and the additional data are available at http://dnaasm.sourceforge.net. (C) 2014 Nalecz Institute of Biocybernetics and Biomedical Engineering. Published by Elsevier Sp. z o.o.. All rights reserved.