AlignWise: a tool for identifying protein-coding sequence and correcting frame-shifts.

AlignWise: a tool for identifying protein-coding sequence and correcting frame-shifts.
复制标题

Alignwise:一种用于识别蛋白质编码序列和校正框架切换的工具。

DOI:
10.1186/s12859-015-0813-8
复制
发表时间:
2015-11-09
期刊:
影响因子:
3
通讯作者:
Loose M
Loose M
中科院分区:
生物学4区
文献类型:
--
作者:
Evans T;Loose M

文献摘要

被引文献

相似文献

从没有参考基因组序列的物种中鉴定蛋白质编码基因可能会因测序错误的存在而变得复杂,特别是插入和缺失。许多工具能够纠正错误的移码组装转录本内是可用的,但往往不报告回DNA序列所需的后续系统发育分析。在这些算法中,Genewise算法是最有效的。然而,它需要一个同源包装器以这种方式使用,在这里,我们证明它完美地纠正帧移只有60%的时间。因此,我们创建了AlignWise,这是一种将Genewise与我们自己的基于同源性的方法AlignFS相结合的工具,用于识别蛋白质编码区并纠正错误的移码,适用于随后的系统发育分析。我们将AlignWise与其他开放式阅读框架查找软件进行了比较,并证明AlignFS算法在纠正订单内的移码方面比Genewise更准确。我们表明,AlignWise在更高的进化距离上提供了最大的准确性,分别优于AlignFS和Genewise。AlignWise为每个转录本生成单个ORF,并高准确性地识别和纠正移码。因此,它非常适合在没有参考基因组的情况下分析新的转录组组装和EST序列。
Identifying protein-coding genes from species without a reference genome sequence can be complicated by the presence of sequencing errors, particularly insertions and deletions. A number of tools capable of correcting erroneous frame-shifts within assembled transcripts are available but often do not report back DNA sequences required for subsequent phylogenetic analysis. Amongst those that do, the Genewise algorithm is the most effective. However, it requires a homology wrapper to be used in this way, and here we demonstrate it perfectly corrects frame-shifts only 60 % of the time. We therefore created AlignWise, a tool that combines Genewise with our own homology-based method, AlignFS, to identify protein-coding regions and correct erroneous frame-shifts, suitable for subsequent phylogenetic analysis. We compared AlignWise against other open reading frame finding software and demonstrate that the AlignFS algorithm is more accurate than Genewise at correcting frame-shifts within an order. We show that AlignWise provides the greatest accuracy at higher evolutionary distances, out-performing both AlignFS and Genewise individually. AlignWise produces a single ORF per transcript and identifies and corrects frame-shifts with high accuracy. It is therefore well suited for analysing novel transcriptome assemblies and EST sequences in the absence of a reference genome.