MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform

MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform
复制标题

DOI:
10.1093/nar/gkf436
复制
发表时间:
2002-07-15
影响因子:
14.9
通讯作者:
Miyata, T
Miyata, T
中科院分区:
生物学2区
文献类型:
--
作者:
Katoh, K;Misawa, K;Miyata, T

文献摘要

被引文献

相似文献

开发了一个多序列比对程序MAFFT。与现有方法相比,CPU时间大大减少。MAFFT包括两种新技术。(i)同源区域通过快速傅立叶变换(FFT)快速识别,其中氨基酸序列被转换成由每个氨基酸残基的体积和极性值组成的序列。(ii)我们提出了一个简化的评分系统,即使对于具有大的插入或延伸的序列以及相似长度的远亲序列,该系统也能很好地减少CPU时间并提高比对的准确性。在MAFFT中实现了两种不同的算法:渐进法(FFT-NS-2)和迭代精化法(FFT-NS-i)。通过计算机仿真和基准测试比较了FFT-NS-2和FFT-NS-i的性能,与CLUSTALW相比,FFT-NS-2的CPU时间大大减少,精度相当。当输入序列的数量超过60时,FFT-NS-i比T-COFFEE快100倍以上,而不牺牲精度。
A multiple sequence alignment program, MAFFT, has been developed. The CPU time is drastically reduced as compared with existing methods. MAFFT includes two novel techniques. (i) Homologous regions are rapidly identified by the fast Fourier transform (FFT), in which an amino acid sequence is converted to a sequence composed of volume and polarity values of each amino acid residue. (ii) We propose a simplified scoring system that performs well for reducing CPU time and increasing the accuracy of alignments even for sequences having large insertions or extensions as well as distantly related sequences of similar length. Two different heuristics, the progressive method (FFT-NS-2) and the iterative refinement method (FFT-NS-i), are implemented in MAFFT. The performances of FFT-NS-2 and FFT-NS-i were compared with other methods by computer simulations and benchmark tests; the CPU time of FFT-NS-2 is drastically reduced as compared with CLUSTALW with comparable accuracy. FFT-NS-i is over 100 times faster than T-COFFEE, when the number of input sequences exceeds 60, without sacrificing the accuracy.