A universal framework for regulatory element discovery across all Genomes and data types

A universal framework for regulatory element discovery across all Genomes and data types
复制标题

DOI:
10.1016/j.molcel.2007.09.027
复制
发表时间:
2007-10-26
期刊:
影响因子:
16
通讯作者:
Tavazoie, Saeed
Tavazoie, Saeed
中科院分区:
生物学1区
文献类型:
--
作者:
Elemento, Olivier;Slonim, Noam;Tavazoie, Saeed

文献摘要

被引文献

相似文献

破译非编码调控基因组已被证明是一项艰巨的挑战。尽管有大量可用的基因表达数据,但目前还没有广泛适用的方法来表征形成丰富的潜在动态的调控元件。我们提出了一个检测此类调控 DNA 和 RNA 基序的通用框架,该框架依赖于直接评估序列和基因表达测量之间的相互信息。我们的方法对背景序列模型和元素影响基因表达的机制做出了最小的假设。这提供了一个跨所有数据类型和基因组的多功能基序发现框架,具有卓越的灵敏度和接近零的假阳性率。从酵母到人类的应用揭示了假定的和已确定的转录因子结合和 miRNA 靶位点,揭示了其空间配置的丰富多样性、DNA 和 RNA 基序的普遍共现、基序回避的上下文依赖选择以及转录后过程对真核转录组的强烈影响。
Deciphering the noncoding regulatory genome has proved a formidable challenge. Despite the wealth of available gene expression data, there currently exists no broadly applicable method for characterizing the regulatory elements that shape the rich underlying dynamics. We present a general framework for detecting such regulatory DNA and RNA motifs that relies on directly assessing the mutual information between sequence and gene expression measurements. Our approach makes minimal assumptions about the background sequence model and the mechanisms by which elements affect gene expression. This provides a versatile motif discovery framework, across all data types and genomes, with exceptional sensitivity and near-zero false-positive rates. Applications from yeast to human uncover putative and established transcription -factor binding and miRNA target sites, revealing rich diversity in their spatial configurations, pervasive cooccurrences of DNA and RNA motifs, context dependent selection for motif avoidance, and the strong impact of post transcriptional processes on eukaryotic transcriptomes.