MUSCLE: multiple sequence alignment with high accuracy and high throughput

MUSCLE: multiple sequence alignment with high accuracy and high throughput
复制标题

DOI:
10.1093/nar/gkh340
复制
发表时间:
2004-03-01
影响因子:
14.9
通讯作者:
Edgar, RC
Edgar, RC
中科院分区:
生物学2区
文献类型:
--
作者:
Edgar, RC

文献摘要

被引文献

相似文献

我们描述肌肉,一个新的计算机程序,用于创建蛋白质序列的多重比对。该算法的元素包括使用kmer计数的快速距离估计、使用我们称为log期望分数的新轮廓函数的渐进对齐以及使用树相关限制分区的细化。MUSCLE的速度和准确性进行了比较,T-Coffee,MAFFT和CLUSTALW的四个测试集的参考比对:BAliBASE,SABmark,SMART和一个新的基准,PREFAB。肌肉达到最高,或联合最高,排名准确性对每一个这些集。在没有细化的情况下,MUSCLE实现了与T-Coffee和MAFFT在统计上无法区分的平均准确度,并且是针对大量序列的测试方法中最快的,在当前台式计算机上在7分钟内对齐平均长度为350的5000个序列。MUSCLE程序、源代码和PREFAB测试数据可在http://www.drive5上免费获得。com/muscle.
We describe MUSCLE, a new computer program for creating multiple alignments of protein sequences. Elements of the algorithm include fast distance estimation using kmer counting, progressive alignment using a new profile function we call the log-expectation score, and refinement using tree-dependent restricted partitioning. The speed and accuracy of MUSCLE are compared with T-Coffee, MAFFT and CLUSTALW on four test sets of reference alignments: BAliBASE, SABmark, SMART and a new benchmark, PREFAB. MUSCLE achieves the highest, or joint highest, rank in accuracy on each of these sets. Without refinement, MUSCLE achieves average accuracy statistically indistinguishable from T-Coffee and MAFFT, and is the fastest of the tested methods for large numbers of sequences, aligning 5000 sequences of average length 350 in 7 min on a current desktop computer. The MUSCLE program, source code and PREFAB test data are freely available at http://www.drive5. com/muscle.