Building a dictionary for genomes: Identification of presumptive regulatory sites by statistical analysis

Building a dictionary for genomes: Identification of presumptive regulatory sites by statistical analysis
复制标题

DOI:
10.1073/pnas.180265397
复制
发表时间:
2000-08-29
影响因子:
11.1
通讯作者:
Siggia, ED
Siggia, ED
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Bussemaker, HJ;Li, H;Siggia, ED

文献摘要

被引文献

相似文献

所有基因的完整基因组序列和mRNA表达数据的可用性为鉴定控制基因表达的DNA序列基序创造了新的机遇和挑战。一个算法,“MobyDick”,提出了一组DNA序列分解成最可能的字典的图案或单词。这种方法适用于任何一组DNA序列:例如,基因组中的所有上游区域或在某些条件下表达的所有基因。单词的识别是基于概率分割模型,其中较长单词的重要性是从各种长度的较短单词的频率推导出来的,从而消除了对单独的一组参考数据来定义概率的需要。我们为酵母基因组中的6,000个上游调控区建立了一个包含1,200个单词的字典; 500个最重要的单词(有些在所有上游区域中只有10个拷贝)匹配443个实验确定的位点中的114个(显著性水平为18个标准差)。当分析所有的基因作为一个组在孢子形成过程中上调,我们发现许多图案,除了少数先前通过分析单独的表达亚群的亚群确定。应用MobyDick的一般阻遏Tup 1被删除时去阻遏的基因,我们发现已知的,以及推定的结合位点,其监管伙伴。
The availability of complete genome sequences and mRNA expression data for all genes creates new opportunities and challenges for identifying DNA sequence motifs that control gene expression. An algorithm, "MobyDick," is presented that decomposes a set of DNA sequences into the most probable dictionary of motifs or words. This method is applicable to any set of DNA sequences: for example, all upstream regions in a genome or all genes expressed under certain conditions. Identification of words is based on a probabilistic segmentation model in which the significance of longer words is deduced from the frequency of shorter ones of various lengths, eliminating the need for a separate set of reference data to define probabilities. We have built a dictionary with 1,200 words for the 6,000 upstream regulatory regions in the yeast genome; the 500 most significant words (some with as few as 10 copies in all of the upstream regions) match 114 of 443 experimentally determined sites (a significance level of 18 standard deviations). When analyzing all of the genes up-regulated during sporulation as a group, we find many motifs in addition to the few previously identified by analyzing the subclusters individually to the expression subclusters. Applying MobyDick to the genes derepressed when the general repressor Tup1 is deleted, we find known as well as putative binding sites for its regulatory partners.