Mining for Putative Regulatory Elements in the Yeast Genome Using Gene Expression Data

Mining for Putative Regulatory Elements in the Yeast Genome Using Gene Expression Data
复制标题

DOI:
--
复制
发表时间:
2000-08
期刊:
Proceedings. International Conference on Intelligent Systems for Molecular Biology
影响因子:
--
通讯作者:
J. Vilo;A. Brazma;I. Jonassen;A. Robinson;E. Ukkonen
J. Vilo;A. Brazma;I. Jonassen;A. Robinson;E. Ukkonen
中科院分区:
其他
文献类型:
--
作者:
J. Vilo;A. Brazma;I. Jonassen;A. Robinson;E. Ukkonen

文献摘要

被引文献

相似文献

我们已经开发了一组方法和工具,用于自动发现基因组序列中推定的调节信号。分析管道包括基因表达数据聚类,从上游序列的基因序列发现的序列模式发现,一个对模式显着性阈值限制检测的控制实验,选择有趣的模式,这些模式的分组,以简洁的形式代表模式组并评估模式组发现针对监管信号现有数据库的推定信号。模式发现在计算上是最昂贵,最关键的步骤。我们的工具对不受限制的长度的先验未知统计上的显着序列模式进行了快速详尽的搜索。相对于一组背景序列,确定了每个群集中一组序列的统计显着性,从而允许检测每个群集特定的微妙调节信号。通过相互相似性将大量的重要模式降低到少数组。这些组的自动得出的共识模式以人为研究者的全面方式代表结果。我们已经对酿酒酵母进行了系统分析。我们创建了大量的表达数据的独立聚类,同时评估了每个群集的“好处”。对于以这种方式获取的52,000多个簇中的每个簇中的每个群集中的每个簇中,我们发现了各个基因上游序列中的重要模式。我们通过形式标准选择了近1,500个重要模式,并将它们与SCPD数据库中实验映射的转录因子结合位点进行了匹配。我们将1,500个模式聚集在62个组中,我们得出了自动对齐和共识模式。在这62个组中,有48个模式在SCPD数据库中具有匹配位点。
We have developed a set of methods and tools for automatic discovery of putative regulatory signals in genome sequences. The analysis pipeline consists of gene expression data clustering, sequence pattern discovery from upstream sequences of genes, a control experiment for pattern significance threshold limit detection, selection of interesting patterns, grouping of these patterns, representing the pattern groups in a concise form and evaluating the discovered putative signals against existing databases of regulatory signals. The pattern discovery is computationally the most expensive and crucial step. Our tool performs a rapid exhaustive search for a priori unknown statistically significant sequence patterns of unrestricted length. The statistical significance is determined for a set of sequences in each cluster with respect to a set of background sequences allowing the detection of subtle regulatory signals specific for each cluster. The potentially large number of significant patterns is reduced to a small number of groups by clustering them by mutual similarity. Automatically derived consensus patterns of these groups represent the results in a comprehensive way for a human investigator. We have performed a systematic analysis for the yeast Saccharomyces cerevisiae. We created a large number of independent clusterings of expression data simultaneously assessing the "goodness" of each cluster. For each of the over 52,000 clusters acquired in this way we discovered significant patterns in the upstream sequences of respective genes. We selected nearly 1,500 significant patterns by formal criteria and matched them against the experimentally mapped transcription factor binding sites in the SCPD database. We clustered the 1,500 patterns to 62 groups for which we derived automatically alignments and consensus patterns. Of these 62 groups 48 had patterns that have matching sites in SCPD database.