Refined repetitive sequence searches utilizing a fast hash function and cross species information retrievals

Refined repetitive sequence searches utilizing a fast hash function and cross species information retrievals
复制标题

DOI:
10.1186/1471-2105-6-111
复制
发表时间:
2005-05-03
期刊:
影响因子:
3
通讯作者:
Shyu, CR
Shyu, CR
中科院分区:
生物学4区
文献类型:
--
作者:
Reneker, J;Shyu, CR

文献摘要

被引文献

相似文献

背景:寻找小串联/分散重复DNA序列简化了许多生物医学研究过程。例如,酵母中的全基因组阵列分析揭示了22个受PHO调控的基因。除了一个之外,它们的启动子区都含有两个核心Pho 4p结合位点CACGTG和CACGTT中的至少一个。在人类中,微卫星在一些罕见的神经退行性疾病中发挥作用,如脊髓小脑共济失调1型(SCA 1)。SCA 1是一种遗传性神经退行性疾病,由基因编码序列中CAG重复序列扩增引起。在细菌病原体中,微卫星被认为可以调控某些毒力因子的表达。例如,细菌通常通过与毒力决定因子密切相关的相变化产生菌株内多样性。最近对幽门螺杆菌菌株26695和J 99的完整序列的分析已经确定了46个假定的相位可变基因,这两个基因组之间通过它们与同聚体束和二核苷酸重复的关联。生命科学家对研究DNA小序列的功能越来越感兴趣。然而,目前的搜索算法往往会产生数以千计的匹配-其中大部分是不相关的researcher.Results:我们提出了我们的哈希函数,以及我们的搜索算法来定位多个基因组内的小序列的DNA。我们的系统应用信息检索算法来发现重复序列的跨物种保护的知识。我们讨论我们的基因本体(GO)数据库纳入这些算法。我们进行了详尽的时间分析,我们的系统的各种重复序列长度。例如,在49条不同的染色体上搜索3.224 GBases内的8个碱基序列平均需要1.147秒。为了说明检索结果的相关性,我们对酵母Pho 4p结合位点CACGTG和CACGTT进行了添加和不添加注释项的检索。此外,一个跨物种的搜索,以说明如何在基因组数据中的潜在隐藏的相关性,可以快速识别。在一个物种中的发现被用作催化剂,以发现另一个物种中的新事物。这些实验也表明,我们的系统表现良好,同时搜索多个基因组-没有主内存的限制,目前在其他system.Conclusion:我们提出了一个时间效率高的算法来定位小片段的DNA,并同时搜索伴随着序列的注释数据。对短序列的全基因组搜索通常会返回数百个匹配结果。我们的实验表明,随后搜索的注释数据可以细化和集中的结果,为用户。我们的算法在主存要求方面也具有空间效率。源代码可根据要求提供。
Background: Searching for small tandem/disperse repetitive DNA sequences streamlines many biomedical research processes. For instance, whole genomic array analysis in yeast has revealed 22 PHO-regulated genes. The promoter regions of all but one of them contain at least one of the two core Pho4p binding sites, CACGTG and CACGTT. In humans, microsatellites play a role in a number of rare neurodegenerative diseases such as spinocerebellar ataxia type 1 (SCA1). SCA1 is a hereditary neurodegenerative disease caused by an expanded CAG repeat in the coding sequence of the gene. In bacterial pathogens, microsatellites are proposed to regulate expression of some virulence factors. For example, bacteria commonly generate intra-strain diversity through phase variation which is strongly associated with virulence determinants. A recent analysis of the complete sequences of the Helicobacter pylori strains 26695 and J99 has identified 46 putative phase-variable genes among the two genomes through their association with homopolymeric tracts and dinucleotide repeats. Life scientists are increasingly interested in studying the function of small sequences of DNA. However, current search algorithms often generate thousands of matches - most of which are irrelevant to the researcher.Results: We present our hash function as well as our search algorithm to locate small sequences of DNA within multiple genomes. Our system applies information retrieval algorithms to discover knowledge of cross-species conservation of repeat sequences. We discuss our incorporation of the Gene Ontology ( GO) database into these algorithms. We conduct an exhaustive time analysis of our system for various repetitive sequence lengths. For instance, a search for eight bases of sequence within 3.224 GBases on 49 different chromosomes takes 1.147 seconds on average. To illustrate the relevance of the search results, we conduct a search with and without added annotation terms for the yeast Pho4p binding sites, CACGTG and CACGTT. Also, a cross-species search is presented to illustrate how potential hidden correlations in genomic data can be quickly discerned. The findings in one species are used as a catalyst to discover something new in another species. These experiments also demonstrate that our system performs well while searching multiple genomes - without the main memory constraints present in other systems.Conclusion: We present a time-efficient algorithm to locate small segments of DNA and concurrently to search the annotation data accompanying the sequence. Genome-wide searches for short sequences often return hundreds of hits. Our experiments show that subsequently searching the annotation data can refine and focus the results for the user. Our algorithms are also space-efficient in terms of main memory requirements. Source code is available upon request.