Correcting errors in shotgun sequences

Correcting errors in shotgun sequences
复制标题

DOI:
10.1093/nar/gkg653
复制
发表时间:
2003-08-01
影响因子:
14.9
通讯作者:
Andersson, B
Andersson, B
中科院分区:
生物学2区
文献类型:
--
作者:
Tammi, MT;Arner, E;Andersson, B

文献摘要

被引文献

相似文献

结合重复区域的测序误差会导致shot弹枪测序的主要问题,这主要是由于组装程序未能区分重复副本之间的单个基本差异与错误的基本调用。在本文中,提出了一种新的策略,旨在使用定义的核苷酸位置DNP纠正shot弹枪序列数据中的错误。该方法通过分析由读取及其与其他读取的所有重叠组成的多个比对来区分单基差与测序误差。使用新型模式匹配算法进行多个比对的构建,该算法利用了可以计算出相同长度相似单词的索引之间的对称性。这允许快速构建多个对齐,而不需要以前的序列读取的配对匹配。该方法的C ++实现的结果表明,可以纠正多达99%的测序误差,而多达87%的单个基本差异仍保留,最多可校正后的读取最多包含一个错误。结果还表明,该方法的表现优于Euler组装程序中使用的误差校正方法。作者可以免费获得原型软件,以供学术使用。
Sequencing errors in combination with repeated regions cause major problems in shotgun sequencing, mainly due to the failure of assembly programs to distinguish single base differences between repeat copies from erroneous base calls. In this paper, a new strategy designed to correct errors in shotgun sequence data using defined nucleotide positions, DNPs, is presented. The method distinguishes single base differences from sequencing errors by analyzing multiple alignments consisting of a read and all its overlaps with other reads. The construction of multiple alignments is performed using a novel pattern matching algorithm, which takes advantage of the symmetry between indices that can be computed for similar words of the same length. This allows for rapid construction of multiple alignments, with no previous pair-wise matching of sequence reads required. Results from a C++ implementation of this method show that up to 99% of sequencing errors can be corrected, while up to 87% of the single base differences remain and up to 80% of the corrected reads contain at most one error. The results also show that the method outperforms the error correction method used in the EULER assembler. The prototype software, MisEd, is freely available from the authors for academic use.