Tandem repeats finder: a program to analyze DNA sequences

Tandem repeats finder: a program to analyze DNA sequences
复制标题

DOI:
10.1093/nar/27.2.573
复制
发表时间:
1999-01-15
影响因子:
14.9
通讯作者:
Benson, G
Benson, G
中科院分区:
生物学2区
文献类型:
--
作者:
Benson, G

文献摘要

被引文献

相似文献

DNA中的串联重复序列是核苷酸模式的两个或更多个连续的近似拷贝。串联重复序列已被证明会导致人类疾病,可能发挥各种调控和进化作用,是重要的实验室和分析工具。关于串联重复序列的模式大小、拷贝数、突变历史等的广泛知识受到无法在基因组序列数据中容易地检测它们的限制。在本文中,我们提出了一个新的算法,发现串联重复,它的工作原理,而不需要指定的模式或模式的大小。我们通过相邻模式拷贝之间的插入缺失的百分比同一性和频率对串联重复序列进行建模,并使用基于统计的识别标准。我们证明了该算法的速度和能力,检测串联重复序列,已经经历了广泛的突变变化,通过分析四个序列:人类共济失调蛋白基因,人类β T细胞受体基因座序列和两个酵母染色体。这些序列的大小范围从3kb到700kb,已经建立了万维网服务器接口c3.biomath.mssm.edu/trf.html用于程序的自动化使用。
A tandem repeat in DNA is two or more contiguous, approximate copies of a pattern of nucleotides. Tandem repeats have been shown to cause human disease, may play a variety of regulatory and evolutionary roles and are important laboratory and analytic tools. Extensive knowledge about pattern size, copy number, mutational history, etc. for tandem repeats has been limited by the inability to easily detect them in genomic sequence data. In this paper, we present a new algorithm for finding tandem repeats which works without the need to specify either the pattern or pattern size. We model tandem repeats by percent identity and frequency of indels between adjacent pattern copies and use statistically based recognition criteria. We demonstrate the algorithm's speed and its ability to detect tandem repeats that have undergone extensive mutational change by analyzing four sequences: the human frataxin gene, the human beta T cell receptor locus sequence and two yeast chromosomes. These sequences range in size from 3 kb up to 700 kb, A World Wide Web server interface at c3.biomath.mssm.edu/trf.html has been established for automated use of the program.