MISHIMA--a new method for high speed multiple alignment of nucleotide sequences of bacterial genome scale data.

MISHIMA--a new method for high speed multiple alignment of nucleotide sequences of bacterial genome scale data.
复制标题

DOI:
10.1186/1471-2105-11-142
复制
发表时间:
2010-03-18
期刊:
影响因子:
3
通讯作者:
Saitou N
Saitou N
中科院分区:
生物学4区
文献类型:
--
作者:
Kryukov K;Saitou N

文献摘要

参考文献

被引文献

相似文献

大型核苷酸序列数据集正成为越来越普遍的比较对象。几乎每天都有完整的细菌基因组的报道。这为开发新的多序列比对方法带来了挑战。传统的多重对准方法是基于成对对准和/或渐进对准技术。这些方法在序列数量较大和处理基因组规模序列时存在性能问题。我们提出了一种新的多序列比对方法,称为MISHIMA (method for Inferring sequence History In multiple alignment),它不依赖于成对序列比较。提出了一种快速查找所有序列共享的稀有寡核苷酸序列的算法。然后采用分而治之的方法将序列分解成可由外部比对程序独立比对的片段。这些部分比对组合在一起形成原始序列的完整比对。与常用的多重对准方法相比,MISHIMA提供了更好的性能。以幽门螺杆菌(Helicobacter pylori)为例,利用一台PC机,在大约6小时内成功比对了6个细菌物种的全基因组序列(每个约1.7 Mb)。
Large nucleotide sequence datasets are becoming increasingly common objects of comparison. Complete bacterial genomes are reported almost everyday. This creates challenges for developing new multiple sequence alignment methods. Conventional multiple alignment methods are based on pairwise alignment and/or progressive alignment techniques. These approaches have performance problems when the number of sequences is large and when dealing with genome scale sequences. We present a new method of multiple sequence alignment, called MISHIMA (Method for Inferring Sequence History In terms of Multiple Alignment), that does not depend on pairwise sequence comparison. A new algorithm is used to quickly find rare oligonucleotide sequences shared by all sequences. Divide and conquer approach is then applied to break the sequences into fragments that can be aligned independently by an external alignment program. These partial alignments are assembled together to form a complete alignment of the original sequences. MISHIMA provides improved performance compared to the commonly used multiple alignment methods. As an example, six complete genome sequences of bacteria species Helicobacter pylori (about 1.7 Mb each) were successfully aligned in about 6 hours using a single PC.
DOI: 10.1093/nar/24.8.1515
发表时间: 1996-04-15
影响因子: 14.9
作者:
Notredame, C;Higgins, DG
通讯作者: Higgins, DG
DOI: 10.1093/bioinformatics/btg192
发表时间: 2003-08-12
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Li, KB
通讯作者: Li, KB
DOI: 10.1093/nar/16.22.10881
发表时间: 1988-11-25
影响因子: 14.9
作者:
CORPET, F
通讯作者: CORPET, F
DOI: 10.1093/molbev/msg139
发表时间: 2003-08-01
影响因子: 10.7
作者:
Shih, ACC;Li, WH
通讯作者: Li, WH
DOI: 10.1006/jmbi.1996.0679
发表时间: 1996-12-13
影响因子: 5.6
作者:
Gotoh, O
通讯作者: Gotoh, O