Discovering motifs in ranked lists of DNA sequences.

Discovering motifs in ranked lists of DNA sequences.
复制标题

DOI:
10.1371/journal.pcbi.0030039
复制
发表时间:
2007-03-23
影响因子:
4.3
通讯作者:
Yakhini Z
Yakhini Z
中科院分区:
生物学2区
文献类型:
--
作者:
Eden E;Lipson D;Yogev S;Yakhini Z

文献摘要

参考文献

被引文献

相似文献

用于发现与背景组相比在靶组中富集的序列元件的计算方法是分子生物学研究中的基础。一个例子是发现转录因子结合基序,从ChIP芯片(染色质免疫沉淀微阵列)测量推断。序列基序发现中的几个主要挑战仍然需要考虑:(i)需要一个原则性的方法来将数据划分为目标和背景集;(ii)缺乏严格的模型和精确的p值来测量基序富集;(iii)需要一个适当的框架来解释基序的多样性;(iv)在许多现有方法中,即使应用于随机生成的数据,也倾向于报告可能重要的基序。在本文中,我们提出了一个统计框架,发现丰富的序列元素的排名列表,解决这四个问题。我们演示了这个框架的实施在一个软件应用程序中,称为DRIM(发现排名不平衡的图案),它确定的序列序列排列的DNA序列列表中的图案。我们将DRIM应用于ChIP芯片和CpG甲基化数据,并获得了以下结果。(i)在酵母ChIP芯片数据中鉴定50个新的推定转录因子(TF)结合位点。对其中一些基因的生物学功能进行了进一步的研究,以期对酵母中的转录调控网络有新的认识。例如,我们的发现能够阐明TF ARO 80的网络。另一个发现涉及系统的TF结合增强含有CA重复序列。(ii)在人类癌症CpG甲基化数据中发现新的基序。值得注意的是,这些基序中的大多数与促进组蛋白甲基化的Polycomb复合物结合的DNA序列元件相似。因此,我们的研究结果支持了一个模型,其中组蛋白甲基化和CpG甲基化是机械连接。总体而言,我们证明了DRIM软件工具中体现的统计框架对于在从表达和ChIP芯片到CpG甲基化数据的各种应用中识别调控序列元件是非常有效的。DRIM可在http://bioinfo.cs.technion.ac.il/drim上公开获取。分子生物学中许多应用的计算问题是鉴定相对于背景基因组序列组在目标基因组序列组中显著过度表达的短DNA序列模式(基序)。一个实例是靶集合,其含有实验测量为结合的特定转录因子蛋白质的DNA序列,而背景集合含有未结合相同转录因子的序列。靶集合中过度表达的序列基序可以代表被转录因子分子识别的子序列。问题的上述表述的固有限制在于,在许多情况下,数据不能以生物学合理的方式被清楚地划分为不同的目标和背景集合。我们描述了一种用于发现基因组序列列表中的基序的统计框架,所述基序根据生物参数或测量(例如,转录因子与序列结合测量)。我们的方法规避了需要分区的数据到目标和背景集使用任意设置的参数。该框架在一个名为DRIM的软件工具中实现。DRIM的应用导致了酵母中新的推定转录因子结合位点的鉴定,以及人类癌细胞系中CpG甲基化区域中先前未知的基序的发现。
Computational methods for discovery of sequence elements that are enriched in a target set compared with a background set are fundamental in molecular biology research. One example is the discovery of transcription factor binding motifs that are inferred from ChIP–chip (chromatin immuno-precipitation on a microarray) measurements. Several major challenges in sequence motif discovery still require consideration: (i) the need for a principled approach to partitioning the data into target and background sets; (ii) the lack of rigorous models and of an exact p-value for measuring motif enrichment; (iii) the need for an appropriate framework for accounting for motif multiplicity; (iv) the tendency, in many of the existing methods, to report presumably significant motifs even when applied to randomly generated data. In this paper we present a statistical framework for discovering enriched sequence elements in ranked lists that resolves these four issues. We demonstrate the implementation of this framework in a software application, termed DRIM (discovery of rank imbalanced motifs), which identifies sequence motifs in lists of ranked DNA sequences. We applied DRIM to ChIP–chip and CpG methylation data and obtained the following results. (i) Identification of 50 novel putative transcription factor (TF) binding sites in yeast ChIP–chip data. The biological function of some of them was further investigated to gain new insights on transcription regulation networks in yeast. For example, our discoveries enable the elucidation of the network of the TF ARO80. Another finding concerns a systematic TF binding enhancement to sequences containing CA repeats. (ii) Discovery of novel motifs in human cancer CpG methylation data. Remarkably, most of these motifs are similar to DNA sequence elements bound by the Polycomb complex that promotes histone methylation. Our findings thus support a model in which histone methylation and CpG methylation are mechanistically linked. Overall, we demonstrate that the statistical framework embodied in the DRIM software tool is highly effective for identifying regulatory sequence elements in a variety of applications ranging from expression and ChIP–chip to CpG methylation data. DRIM is publicly available at http://bioinfo.cs.technion.ac.il/drim. A computational problem with many applications in molecular biology is to identify short DNA sequence patterns (motifs) that are significantly overrepresented in a target set of genomic sequences relative to a background set of genomic sequences. One example is a target set that contains DNA sequences to which a specific transcription factor protein was experimentally measured as bound while the background set contains sequences to which the same transcription factor was not bound. Overrepresented sequence motifs in the target set may represent a subsequence that is molecularly recognized by the transcription factor. An inherent limitation of the above formulation of the problem lies in the fact that in many cases data cannot be clearly partitioned into distinct target and background sets in a biologically justified manner. We describe a statistical framework for discovering motifs in a list of genomic sequences that are ranked according to a biological parameter or measurement (e.g., transcription factor to sequence binding measurements). Our approach circumvents the need to partition the data into target and background sets using arbitrarily set parameters. The framework is implemented in a software tool called DRIM. The application of DRIM led to the identification of novel putative transcription factor binding sites in yeast and to the discovery of previously unknown motifs in CpG methylation regions in human cancer cell lines.
DOI: 10.1016/j.cell.2005.10.042
发表时间: 2006-01-13
期刊: CELL
影响因子: 64.5
作者:
Hallikas, O;Palin, K;Taipale, J
通讯作者: Taipale, J
DOI: 10.1093/bioinformatics/bti402
发表时间: 2005-06-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Hong, PY;Liu, XS;Wong, WH
通讯作者: Wong, WH
DOI: 10.1186/1471-2105-6-84
发表时间: 2005-04-04
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Friberg, M;von Rohr, P;Gonnet, G
通讯作者: Gonnet, G
DOI: 10.1089/106652700750050943
发表时间: 2000-01-01
影响因子: 1.7
作者:
Ben-Dor, A;Bruhn, L;Yakhini, Z
通讯作者: Yakhini, Z
DOI: 10.1038/84792
发表时间: 2001-02-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Bussemaker, HJ;Li, H;Siggia, ED
通讯作者: Siggia, ED