DIALIGN-TX: greedy and progressive approaches for segment-based multiple sequence alignment

DIALIGN-TX: greedy and progressive approaches for segment-based multiple sequence alignment
复制标题

DOI:
10.1186/1748-7188-3-6
复制
发表时间:
2008-05-27
影响因子:
1
通讯作者:
Morgenstern, Burkhard
Morgenstern, Burkhard
中科院分区:
生物学4区
文献类型:
--
作者:
Subramanian, Amarendran R.;Kaufmann, Michael;Morgenstern, Burkhard

文献摘要

被引文献

相似文献

背景:DIALIGN-T是多重比对程序DIALIGN的重新实现。由于几个算法的改进,它在局部和全局相关的序列集上比以前的DIALIGN版本产生了显著更好的比对。然而,像程序的原始实现一样,DIALIGN-T使用一种直接的贪婪方法来根据局部成对序列相似性组装多个比对。这种贪婪的方法可能容易受到虚假的随机相似性的影响,因此可能导致次优结果。在本文中,我们提出了DIALIGN-TX,它是对DIALIGN-T的实质性改进,它结合了我们以前的贪婪算法和渐进比对方法。结果:我们的新启发式算法在不过度增加CPU时间和内存消耗的情况下,得到了明显更好的比对结果,尤其是在全局相关序列上。新方法基于引导树;为了检测可能的伪序列相似性,它在冲突图上使用顶点覆盖近似。我们对一大组核酸和蛋白质序列进行了基准测试,以确定蛋白质基准,我们分别使用基准数据库BALIBASE 3和数据库IRMBASE 2的更新版本来评估全球和局部相关序列的质量。对于核酸序列的比对,我们使用BRAliBase II进行全球比对,并建立了一个新的局部相关序列数据库DIRM-BASE 1。IRMBase 2和DIRMBase 1是通过在长而不可比对的序列中的随机位置植入高度保守的动机来构建的。结论:在BALIBASE 3上,我们的新程序比以前的程序DIALIGN-T性能要好得多,也优于流行的全局比对程序CLUSTAL W,但它的性能仍然优于MAFFT、MASK和T-CAFEE等专注于全局比对的程序。在IRMBASE 2和DIRM-BASE 1的局部相关测试集上,我们的方法优于所有其他程序,而MAFFT E-INSI是唯一接近DIALIGN-TX性能的方法。
Background: DIALIGN-T is a reimplementation of the multiple-alignment program DIALIGN. Due to several algorithmic improvements, it produces significantly better alignments on locally and globally related sequence sets than previous versions of DIALIGN. However, like the original implementation of the program, DIALIGN-T uses a a straight-forward greedy approach to assemble multiple alignments from local pairwise sequence similarities. Such greedy approaches may be vulnerable to spurious random similarities and can therefore lead to suboptimal results. In this paper, we present DIALIGN-TX, a substantial improvement of DIALIGN-T that combines our previous greedy algorithm with a progressive alignment approach.Results: Our new heuristic produces significantly better alignments, especially on globally related sequences, without increasing the CPU time and memory consumption exceedingly. The new method is based on a guide tree; to detect possible spurious sequence similarities, it employs a vertex-cover approximation on a conflict graph. We performed benchmarking tests on a large set of nucleic acid and protein sequences For protein benchmarks we used the benchmark database BALIBASE 3 and an updated release of the database IRMBASE 2 for assessing the quality on globally and locally related sequences, respectively. For alignment of nucleic acid sequences, we used BRAliBase II for global alignment and a newly developed database of locally related sequences called DIRM-BASE 1. IRMBASE 2 and DIRMBASE 1 are constructed by implanting highly conserved motives at random positions in long unalignable sequences.Conclusion: On BALIBASE 3, our new program performs significantly better than the previous program DIALIGN-T and outperforms the popular global aligner CLUSTAL W, though it is still outperformed by programs that focus on global alignment like MAFFT, MUSCLE and T-COFFEE. On the locally related test sets in IRMBASE 2 and DIRM-BASE 1, our method outperforms all other programs while MAFFT E-INSi is the only method that comes close to the performance of DIALIGN-TX.