transAlign: using amino acids to facilitate the multiple alignment of protein-coding DNA sequences.

transAlign: using amino acids to facilitate the multiple alignment of protein-coding DNA sequences.
复制标题

DOI:
10.1186/1471-2105-6-156
复制
发表时间:
2005-06-22
期刊:
影响因子:
3
通讯作者:
Bininda-Emonds OR
Bininda-Emonds OR
中科院分区:
生物学4区
文献类型:
--
作者:
Bininda-Emonds OR

文献摘要

参考文献

被引文献

相似文献

同源DNA序列的比对对于比较基因组学和系统发育分析是至关重要的。然而,多重比对代表了计算上困难的问题。对于编码蛋白质的DNA序列,比对由DNA序列指定的氨基酸序列而不是DNA序列本身在速度和准确性方面更有利。利用“翻译比对”的概念的许多实施方式在它们需要用户手动翻译DNA序列并进行氨基酸比对的意义上是不完整的。因此,它们不太适合大规模和/或众多DNA数据集的大规模自动比对。transAlign是一个开源Perl脚本,它通过氨基酸翻译来比对蛋白质编码DNA序列,以利用氨基酸比对的上级多重比对能力和速度。它的工作原理是将每个DNA序列翻译成相应的氨基酸序列,将整个矩阵传递给ClustalW进行比对,然后反向翻译所得的氨基酸比对,以获得比对的DNA序列。在翻译步骤中,transAlign根据所需的遗传密码确定每个DNA序列的最佳方向和阅读框架。它还可以检查DNA序列中的明显移码,并可以以三种方式之一处理移码序列(删除,无论如何作为氨基酸对齐,或作为DNA进行轮廓对齐)。正如一组来自哺乳动物六个蛋白质编码基因的比较基准所示,transAlign中实施的策略总是提高蛋白质编码DNA序列比对的速度和通常的表观准确度。transAlign是翻译比对概念的少数几个完整和跨平台实现之一。从执行翻译比对中获得的优势和程序中可用的用户可定义选项套件意味着transAlign非常适合非常大和/或非常多的蛋白质编码DNA数据集的大规模自动比对。然而,该程序提供的良好性能也转化为任何一组蛋白质编码序列的比对。transAlign,包括源代码,可以在http://www.tierzucht.tum.de/Bininda-Emonds/(在“程序”下)免费获得。
Alignments of homologous DNA sequences are crucial for comparative genomics and phylogenetic analysis. However, multiple alignment represents a computationally difficult problem. For protein-coding DNA sequences, it is more advantageous in terms of both speed and accuracy to align the amino-acid sequences specified by the DNA sequences rather than the DNA sequences themselves. Many implementations making use of this concept of "translated alignments" are incomplete in the sense that they require the user to manually translate the DNA sequences and to perform the amino-acid alignment. As such, they are not well suited to large-scale automated alignments of large and/or numerous DNA data sets. transAlign is an open-source Perl script that aligns protein-coding DNA sequences via their amino-acid translations to take advantage of the superior multiple-alignment capabilities and speed of an amino-acid alignment. It operates by translating each DNA sequence into its corresponding amino-acid sequence, passing the entire matrix to ClustalW for alignment, and then back-translating the resulting amino-acid alignment to derive the aligned DNA sequences. In the translation step, transAlign determines the optimal orientation and reading frame for each DNA sequence according to the desired genetic code. It also checks for apparent frame shifts in the DNA sequences and can handle frame-shifted sequences in one of three ways (delete, align as amino acids regardless, or profile align as DNA). As a set of comparative benchmarks derived from six protein-coding genes for mammals shows, the strategy implemented in transAlign always improves the speed and usually the apparent accuracy of the alignment of protein-coding DNA sequences. transAlign represents one of few full and cross-platform implementations of the concept of translated alignments. Both the advantages accruing from performing a translated alignment and the suite of user-definable options available in the program mean that transAlign is ideally suited for large-scale automated alignments of very large and/or very numerous protein-coding DNA data sets. However, the good performance offered by the program also translates to the alignment of any set of protein-coding sequences. transAlign, including the source code, is freely available at http://www.tierzucht.tum.de/Bininda-Emonds/ (under "Programs").
DOI: 10.1073/pnas.89.22.10915
发表时间: 1992-11-15
影响因子: 11.1
作者:
HENIKOFF, S;HENIKOFF, JG
通讯作者: HENIKOFF, JG
DOI: 10.1080/106351501750435121
发表时间: 2001-08-01
期刊: SYSTEMATIC BIOLOGY
影响因子: 6.5
作者:
Posada, D;Crandall, KA
通讯作者: Crandall, KA
DOI: 10.1093/nar/gkg500
发表时间: 2003-07-01
影响因子: 14.9
作者:
Chenna, R;Sugawara, H;Thompson, JD
通讯作者: Thompson, JD
DOI: 10.1093/nar/30.1.169
发表时间: 2002-01-01
影响因子: 14.9
作者:
Wain, HM;Lush, M;Povey, S
通讯作者: Povey, S
DOI: 10.1093/bioinformatics/15.3.211
发表时间: 1999-03-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Morgenstern, B
通讯作者: Morgenstern, B