Genome comparison without alignment using shortest unique substrings.

Genome comparison without alignment using shortest unique substrings.
复制标题

DOI:
10.1186/1471-2105-6-123
复制
发表时间:
2005-05-23
期刊:
影响因子:
3
通讯作者:
Wiehe T
Wiehe T
中科院分区:
生物学4区
文献类型:
--
作者:
Haubold B;Pierstorff N;Möller F;Wiehe T

文献摘要

参考文献

被引文献

相似文献

通过比对进行序列比较是分子生物学的基本工具。在本文中,我们展示了如何在没有比对步骤的情况下有效地完成许多序列比较任务,包括检测独特的基因组区域。我们的核苷酸序列比较程序基于最短的独特子串。这些子串在所分析的序列或序列集中仅出现一次,并且在不失去唯一性的情况下不能进一步减少长度。可以使用广义后缀树来检测此类子字符串。我们发现,在秀丽隐杆线虫、人类和小鼠的常染色体中,最短的独特子串不超过 11 bp。在小鼠和人类中,这些独特的子串明显聚集在已知基因的上游区域。而且,在人类或小鼠的基因组中偶然发现如此短的独特子串的概率极小。给定查询序列的 GC 内容,我们得出最短唯一子串的空分布的分析表达式。此外,我们应用我们的方法快速检测金黄色葡萄球菌菌株 MSSA476 基因组中与其他四种葡萄球菌基因组相比的独特基因组区域。我们结合了一种方法来快速搜索 DNA 序列中最短的唯一子串及其零分布的推导。我们表明,用这种方法可以有效地检测任意基因组样本中的独特区域。相应的程序 shustring (SHortest Unique subSTRING) 和 shulen 是用 C 编写的,可从 获取。
Sequence comparison by alignment is a fundamental tool of molecular biology. In this paper we show how a number of sequence comparison tasks, including the detection of unique genomic regions, can be accomplished efficiently without an alignment step. Our procedure for nucleotide sequence comparison is based on shortest unique substrings. These are substrings which occur only once within the sequence or set of sequences analysed and which cannot be further reduced in length without losing the property of uniqueness. Such substrings can be detected using generalized suffix trees. We find that the shortest unique substrings in Caenorhabditis elegans, human and mouse are no longer than 11 bp in the autosomes of these organisms. In mouse and human these unique substrings are significantly clustered in upstream regions of known genes. Moreover, the probability of finding such short unique substrings in the genomes of human or mouse by chance is extremely small. We derive an analytical expression for the null distribution of shortest unique substrings, given the GC-content of the query sequences. Furthermore, we apply our method to rapidly detect unique genomic regions in the genome of Staphylococcus aureus strain MSSA476 compared to four other staphylococcal genomes. We combine a method to rapidly search for shortest unique substrings in DNA sequences and a derivation of their null distribution. We show that unique regions in an arbitrary sample of genomes can be efficiently detected with this method. The corresponding programs shustring (SHortest Unique subSTRING) and shulen are written in C and available at .
DOI: 10.1126/science.270.5235.397
发表时间: 1995-10-20
期刊: SCIENCE
影响因子: 56.9
作者:
FRASER, CM;GOCAYNE, JD;VENTER, JC
通讯作者: VENTER, JC
DOI: 10.1016/s0140-6736(00)04403-2
发表时间: 2001-04-21
期刊: LANCET
影响因子: 168.9
作者:
Kuroda, M;Ohta, T;Hiramatsu, K
通讯作者: Hiramatsu, K
DOI: 10.1006/jmbi.1990.9999
发表时间: 1990-10-05
影响因子: 5.6
作者:
ALTSCHUL, SF;GISH, W;LIPMAN, DJ
通讯作者: LIPMAN, DJ
DOI: 10.1016/0022-2836(70)90057-4
发表时间: 1970-01-01
影响因子: 5.6
作者:
NEEDLEMAN, SB;WUNSCH, CD
通讯作者: WUNSCH, CD
DOI: 10.1016/s0140-6736(02)08713-5
发表时间: 2002-05-25
期刊: LANCET
影响因子: 168.9
作者:
Baba, T;Takeuchi, F;Hiramatsu, K
通讯作者: Hiramatsu, K