De-novo discovery of differentially abundant transcription factor binding sites including their positional preference.

De-novo discovery of differentially abundant transcription factor binding sites including their positional preference.
复制标题

DOI:
10.1371/journal.pcbi.1001070
复制
发表时间:
2011-02-10
影响因子:
4.3
通讯作者:
Grosse I
Grosse I
中科院分区:
生物学2区
文献类型:
--
作者:
Keilwagen J;Grau J;Paponov IA;Posch S;Strickert M;Grosse I

文献摘要

参考文献

被引文献

相似文献

转录因子是基因调控的主要组成部分,因为它们通过与启动子中的特定结合位点结合来激活或抑制基因表达。通过湿实验室实验获得的靶区域转录因子结合位点的从头发现是计算生物学中的一个具有挑战性的问题,目前尚未完全解决。在这里,我们提出了一种名为 Dispom 的从头基序发现工具,用于寻找差异丰富的转录因子结合位点,模拟结合位点的现有位置偏好并在学习过程中调整基序的长度。通过评估 Dispom,我们发现它的预测性能优于现有的从头基序发现工具,适用于 18 个具有植入结合位点的基准数据集,以及基于微阵列、ChIP 芯片、ChIP-DSL 和 DamID 以及基因本体数据的实验数据的后生动物纲要。最后,我们应用 Dispom 来寻找从拟南芥微阵列数据中提取的生长素响应基因的启动子中差异丰富的结合位点,并且我们发现了一个可以解释为精炼的生长素响应元件的基序,主要位于转录起始位点上游 250 bp 区域。使用生长素响应基因的独立数据集,我们在全基因组预测中发现,与典型的生长素响应元件相比,精炼基序对生长素响应基因更具特异性。一般来说,Dispom 可用于查找任何来源序列中差异丰富的基序。然而,如果所有序列都与某个锚点(例如启动子序列的转录起始位点)对齐,则 Dispom 学到的位置分布特别有用。我们证明,搜索差异丰富的基序和从数据推断位置分布的结合有利于从头发现基序。因此,我们将该工具作为开源 Java 框架 Jstacs 的组件以及 http://www.jstacs.de/index.php/Dispom 上的独立应用程序免费提供。转录因子与基因启动子的结合以及随后的转录增强或抑制是转录基因调控的主要步骤之一。直接或间接湿实验室实验可以识别转录因子可能结合或调节的大致区域。随后,从头基序发现工具可用于检测结合位点的精确位置。许多传统工具关注目标区域中过度表达的基序,而这些基序通常在整个基因组中同样过度表达。相比之下,最近的一些工具侧重于目标区域与对照组相比差异丰富的基序。由于结合位点通常位于距转录起始位点的某个优选距离处,因此最好将此信息纳入从头基序发现中。在这里,我们提出了 Dispom 一种同时学习差异丰富的基序及其位置偏好的新方法,与许多流行的从头基序发现工具相比,它可以更准确地预测结合位点。当将 Dispom 应用于拟南芥生长素响应基因的启动子时,我们发现了一个与典型的生长素响应元件略有不同的结合基序,它表现出强烈的位置偏好,并且对生长素响应基因更具特异性。
Transcription factors are a main component of gene regulation as they activate or repress gene expression by binding to specific binding sites in promoters. The de-novo discovery of transcription factor binding sites in target regions obtained by wet-lab experiments is a challenging problem in computational biology, which has not been fully solved yet. Here, we present a de-novo motif discovery tool called Dispom for finding differentially abundant transcription factor binding sites that models existing positional preferences of binding sites and adjusts the length of the motif in the learning process. Evaluating Dispom, we find that its prediction performance is superior to existing tools for de-novo motif discovery for 18 benchmark data sets with planted binding sites, and for a metazoan compendium based on experimental data from micro-array, ChIP-chip, ChIP-DSL, and DamID as well as Gene Ontology data. Finally, we apply Dispom to find binding sites differentially abundant in promoters of auxin-responsive genes extracted from Arabidopsis thaliana microarray data, and we find a motif that can be interpreted as a refined auxin responsive element predominately positioned in the 250-bp region upstream of the transcription start site. Using an independent data set of auxin-responsive genes, we find in genome-wide predictions that the refined motif is more specific for auxin-responsive genes than the canonical auxin-responsive element. In general, Dispom can be used to find differentially abundant motifs in sequences of any origin. However, the positional distribution learned by Dispom is especially beneficial if all sequences are aligned to some anchor point like the transcription start site in case of promoter sequences. We demonstrate that the combination of searching for differentially abundant motifs and inferring a position distribution from the data is beneficial for de-novo motif discovery. Hence, we make the tool freely available as a component of the open-source Java framework Jstacs and as a stand-alone application at http://www.jstacs.de/index.php/Dispom. Binding of transcription factors to promoters of genes, and subsequent enhancement or repression of transcription, is one of the main steps of transcriptional gene regulation. Direct or indirect wet-lab experiments allow the identification of approximate regions potentially bound or regulated by a transcription factor. Subsequently, de-novo motif discovery tools can be used for detecting the precise positions of binding sites. Many traditional tools focus on motifs over-represented in the target regions, which often turn out to be similarly over-represented in the entire genome. In contrast, several recent tools focus on differentially abundant motifs in target regions compared to a control set. As binding sites are often located at some preferred distance to the transcription start site, it is favorable to include this information into de-novo motif discovery. Here, we present Dispom a novel approach for learning differentially abundant motifs and their positional preferences simultaneously, which predicts binding sites with increased accuracy compared to many popular de-novo motif discovery tools. When applying Dispom to promoters of auxin-responsive genes of Arabidopsis thaliana, we find a binding motif slightly different from the canonical auxin-response element, which exhibits a strong positional preference and which is considerably more specific to auxin-responsive genes.
TransFac及其模块移植:真核生物中的转录基因调节。
DOI: 10.1093/nar/gkj143
发表时间: 2006-01-01
影响因子: 14.9
作者:
Matys V;Kel-Margoulis OV;Fricke E;Liebich I;Land S;Barre-Dirrie A;Reuter I;Chekmenev D;Krull M;Hornischer K;Voss N;Stegmaier P;Lewicki-Potapov B;Saxel H;Kel AE;Wingender E
通讯作者: Wingender E
DOI: 10.1186/1471-2105-9-262
发表时间: 2008-06-04
期刊: BMC bioinformatics
影响因子: 3
作者:
Kim NK;Tharakaraman K;Mariño-Ramírez L;Spouge JL
通讯作者: Spouge JL
DOI: 10.1006/abio.1997.2231
发表时间: 1997-08-01
影响因子: 2.9
作者:
Benotmane, AM;Hoylaerts, MF;Belayew, A
通讯作者: Belayew, A
DOI: 10.1007/s00425-004-1206-9
发表时间: 2004-05-01
期刊: PLANTA
影响因子: 4.3
作者:
Mönke, G;Altschmied, L;Conrad, U
通讯作者: Conrad, U
DOI: 10.1007/11564096_12
发表时间: 2005-01-01
期刊: MACHINE LEARNING: ECML 2005, PROCEEDINGS
影响因子: --
作者:
Cerquides, J;de Màntaras, RL
通讯作者: de Màntaras, RL