Instability in progressive multiple sequence alignment algorithms.

Instability in progressive multiple sequence alignment algorithms.
复制标题

DOI:
10.1186/s13015-015-0057-1
复制
发表时间:
2015
期刊:
Algorithms for molecular biology : AMB
影响因子:
--
通讯作者:
Higgins DG
Higgins DG
中科院分区:
其他
文献类型:
--
作者:
Boyce K;Sievers F;Higgins DG

文献摘要

被引文献

相似文献

渐进比对是用于比对大量序列的标准方法。与所有启发式算法一样,这涉及到对齐精度和计算时间之间的权衡。我们检查这种权衡并发现,由于该方法的早期步骤中的信息丢失,由最常见的多序列比对程序生成的比对本质上是不稳定的,并且简单地颠倒输入文件中的序列的顺序将导致生成不同的比对。虽然这种影响在序列数量较多时更加明显,但在一百个序列的数量级的数据集中也可以看到这种影响。我们还概述了确定数据集中序列数量的方法,超过这些序列,不稳定的可能性将变得更加明显。这对大规模多序列比对算法的设计者和这些比对的用户都有重大的影响。
Progressive alignment is the standard approach used to align large numbers of sequences. As with all heuristics, this involves a tradeoff between alignment accuracy and computation time. We examine this tradeoff and find that, because of a loss of information in the early steps of the approach, the alignments generated by the most common multiple sequence alignment programs are inherently unstable, and simply reversing the order of the sequences in the input file will cause a different alignment to be generated. Although this effect is more obvious with larger numbers of sequences, it can also be seen with data sets in the order of one hundred sequences. We also outline the means to determine the number of sequences in a data set beyond which the probability of instability will become more pronounced. This has major ramifications for both the designers of large-scale multiple sequence alignment algorithms, and for the users of these alignments.