Correcting sequencing errors in DNA coding regions using a dynamic programming approach.

Correcting sequencing errors in DNA coding regions using a dynamic programming approach.
复制标题

使用动态编程方法纠正 DNA 编码区的测序错误。

DOI:
10.1093/bioinformatics/11.2.117
复制
发表时间:
1995
期刊:
Computer applications in the biosciences : CABIOS
影响因子:
--
通讯作者:
Uberbacher,EC
Uberbacher,EC
中科院分区:
--
文献类型:
--
作者:
Xu,Y;Mural,RJ;Uberbacher,EC

文献摘要

相似文献

本文提出了一种检测和“纠正”DNA编码区序列错误的算法。涉及的测序错误的类型是DNA碱基的插入和缺失(Indels)。我们的目标是提供一种能力,使单次通过或低冗余序列数据更具信息量,减少用于基因鉴定和表征目的的高冗余测序的需要。这将提高测序效率并降低基因组测序成本。该算法通过发现假设编码区内的统计偏好阅读框中的变化来检测测序错误,然后在感知的阅读框过渡点插入多个‘中性’碱基,以使假设的外显子候选帧保持一致。我们已经实现了该算法作为GRAIL DNA序列分析系统的前端子系统,以构建一个容错能力很强的版本,并打算将其作为进一步开发测序纠错技术的试验台。初步的测试结果表明了该算法的有效性,同时也揭示了它的一些不足之处,为进一步改进提供了可能的方向。在由68个人类DNA序列组成的测试集中,在编码区随机生成1%的INDELs,该算法检测并纠正了76%的INDELs。Indel的位置与预测的位置之间的平均距离为9.4个碱基。有了这个子系统后,GRAIL正确预测了89%的编码消息,其中10%的错误消息出现在‘更正’序列上,相比之下,使用标准GRAIL II方法(1.2版),正确预测的编码消息的比例为69%,错误预测的消息比例为11%。该方法使用动态规划算法,在时间和空间上与输入序列的大小成线性关系。
This paper presents an algorithm for detecting and ‘correcting’ sequencing errors that occur in DNA coding regions. The types of sequencing errors addressed are insertions and deletions (indels) of DNA bases. The goal is to provide a capability which makes single-pass or low-redundancy sequence data more informative, reducing the need for high-redundancy sequencing for gene identification and characterization purposes. This would permit improved sequencing efficiency and reduce genome sequencing costs. The algorithm detects sequencing errors by discovering changes in the statistically preferred reading frame within a putative coding region and then inserts a number of ‘neutral’ bases at a perceived reading frame transition point to make the putative exon candidate frame consistent. We have implemented the algorithm as a front-end subsystem of the GRAIL DNA sequence analysis system to construct a version which is very error tolerant and also intend to use this as a testbed for further development of sequencing error-correction technology. Preliminary test results have shown the usefulness of this algorithm and also exhibited some of its weakness, providing possible directions for further improvement. On a test set consisting of 68 human DNA sequences with 1% randomly generated indels in coding regions, the algorithm detected and corrected 76% of the indels. The average distance between the position of an indel and the predicted one was 9.4 bases. With this subsystem in place, GRAIL correctly predicted 89% of the coding messages with 10% false message on the ‘corrected’ sequences, compared to 69% correctly predicted coding messages and 11% falsely predicted messages on the ‘corrupted’ sequences using standard GRAIL II method (version 1.2). The method uses a dynamic programming algorithm, and runs in time and space linear to the size of the input sequence.