Fast algorithms for large-scale genome alignment and comparison

Fast algorithms for large-scale genome alignment and comparison
复制标题

DOI:
10.1093/nar/30.11.2478
复制
发表时间:
2002-06-01
影响因子:
14.9
通讯作者:
Salzberg, SL
Salzberg, SL
中科院分区:
生物学2区
文献类型:
--
作者:
Delcher, AL;Phillippy, A;Salzberg, SL

文献摘要

被引文献

相似文献

我们描述了一种后缀树算法,该算法可以使真核和原核生物的整个基因组序列与计算机时间和记忆的最小使用。新系统木乃伊2的运行速度快三倍,同时使用三分之一的内存,与原始的木乃伊系统一样多。它已成功地将整个人类和小鼠基因组对齐,并将许多较小的真核生物和原核基因组对齐。一个新的模块允许对多个DNA序列片段的比对,这在比较不完整的基因组序列中被证明是有价值的。我们还描述了一种通过检测蛋白质序列同源性来对齐更遥远的基因组的方法。在翻译所有六个读取帧中的序列后,将所有匹配的蛋白质序列提取,然后将簇放在一起匹配之后,这将使两个基因组对齐两个基因组。该方法已应用于不完整和完整的基因组序列,以检测保守的同义区域,其中一种有生物体的多种蛋白在另一种生物体中以相同的顺序和方向发现。作者可以免费提供系统代码。
We describe a suffix-tree algorithm that can align the entire genome sequences of eukaryotic and prokaryotic organisms with minimal use of computer time and memory. The new system, MUMmer 2, runs three times faster while using one-third as much memory as the original MUMmer system. It has been used successfully to align the entire human and mouse genomes to each other, and to align numerous smaller eukaryotic and prokaryotic genomes. A new module permits the alignment of multiple DNA sequence fragments, which has proven valuable in the comparison of incomplete genome sequences. We also describe a method to align more distantly related genomes by detecting protein sequence homology. This extension to MUMmer aligns two genomes after translating the sequence in all six reading frames, extracts all matching protein sequences and then clusters together matches. This method has been applied to both incomplete and complete genome sequences in order to detect regions of conserved synteny, in which multiple proteins from one organism are found in the same order and orientation in another. The system code is being made freely available by the authors.