Motif Enrichment Analysis: a unified framework and an evaluation on ChIP data

Motif Enrichment Analysis: a unified framework and an evaluation on ChIP data
复制标题

DOI:
10.1186/1471-2105-11-165
复制
发表时间:
2010-04-01
期刊:
影响因子:
3
通讯作者:
Bailey, Timothy L.
Bailey, Timothy L.
中科院分区:
生物学4区
文献类型:
--
作者:
McLeay, Robert C.;Bailey, Timothy L.

文献摘要

被引文献

相似文献

背景:分子生物学的一个主要目标是确定控制基因转录的机制。基序富集分析(MEA)试图通过检测基因调控区中已知结合基序的富集来确定哪些DNA结合转录因子控制一组基因的转录。通常,生物学家指定一组被认为是共调节的基因和已知的转录因子DNA结合模型库,MEA确定哪些因子(如果有的话)可能是基因的直接调节因子。由于已知DNA结合模型的因子数量由于高通量技术而迅速增加,MEA变得越来越有用。在本文中,我们探讨如何使MEA适用于更多的设置,并评估一些MEA的approaches.Results的有效性:我们首先定义了一个数学框架的基序富集分析,放宽了生物学家输入一组选定的基因的要求。相反,输入由所有调节区域组成,每个区域都标记有生物信号的水平。然后,我们定义并实现了一些基序富集分析方法。其中一些方法需要用户指定的信号阈值,一些以数据驱动的方式确定最佳阈值,我们的两种方法是无阈值的。我们评估这些方法,沿着与现有的两种方法(三叶草和PASTAA),使用酵母ChIP芯片数据。我们的新的基于线性回归的无阈值方法在我们的评估中表现最好,其次是数据驱动的PASTAA算法。如果用户指定的阈值被最佳选择,则三叶草算法的性能与PASTAA一样好。基于三种统计检验(Fisher精确检验、秩和检验和多重超几何检验)的数据驱动方法表现不佳,即使在阈值选择最佳时也是如此。这些方法(和三叶草)执行甚至更糟时,不受限制的数据驱动的阈值determinations.Conclusions:我们的新的,无阈值线性回归方法以及ChIP芯片数据。使用数据驱动阈值确定的方法可能执行得很差,除非阈值的范围是先验限制的。然而,PASTAA中实施的限制似乎是精心选择的。我们的新算法-AME(基序富集分析)-可在http://bioinformatics.org.au/ame/上获得。
Background: A major goal of molecular biology is determining the mechanisms that control the transcription of genes. Motif Enrichment Analysis (MEA) seeks to determine which DNA-binding transcription factors control the transcription of a set of genes by detecting enrichment of known binding motifs in the genes' regulatory regions. Typically, the biologist specifies a set of genes believed to be co-regulated and a library of known DNA-binding models for transcription factors, and MEA determines which (if any) of the factors may be direct regulators of the genes. Since the number of factors with known DNA-binding models is rapidly increasing as a result of high-throughput technologies, MEA is becoming increasingly useful. In this paper, we explore ways to make MEA applicable in more settings, and evaluate the efficacy of a number of MEA approaches.Results: We first define a mathematical framework for Motif Enrichment Analysis that relaxes the requirement that the biologist input a selected set of genes. Instead, the input consists of all regulatory regions, each labeled with the level of a biological signal. We then define and implement a number of motif enrichment analysis methods. Some of these methods require a user-specified signal threshold, some identify an optimum threshold in a data-driven way and two of our methods are threshold-free. We evaluate these methods, along with two existing methods (Clover and PASTAA), using yeast ChIP-chip data. Our novel threshold-free method based on linear regression performs best in our evaluation, followed by the data-driven PASTAA algorithm. The Clover algorithm performs as well as PASTAA if the user-specified threshold is chosen optimally. Data-driven methods based on three statistical tests-Fisher Exact Test, rank-sum test, and multi-hypergeometric test-perform poorly, even when the threshold is chosen optimally. These methods (and Clover) perform even worse when unrestricted data-driven threshold determination is used.Conclusions: Our novel, threshold-free linear regression method works well on ChIP-chip data. Methods using data-driven threshold determination can perform poorly unless the range of thresholds is limited a priori. The limits implemented in PASTAA, however, appear to be well-chosen. Our novel algorithms-AME (Analysis of Motif Enrichment)-are available at http://bioinformatics.org.au/ame/.