A feature-based approach to modeling protein-DNA interactions.

A feature-based approach to modeling protein-DNA interactions.
复制标题

DOI:
10.1371/journal.pcbi.1000154
复制
发表时间:
2008-08-22
影响因子:
4.3
通讯作者:
Segal E
Segal E
中科院分区:
生物学2区
文献类型:
--
作者:
Sharon E;Lubliner S;Segal E

文献摘要

参考文献

被引文献

相似文献

转录因子(TF)与其DNA靶位点的结合是一种基本的调控相互作用。用于表示TF结合特异性的最常见的模型是位置特异性评分矩阵(PSSM),其假设结合位置之间的独立性。然而,在许多情况下,这种简化的假设并不成立。在这里,我们提出了特征基序模型(FRENT),一种新的概率方法建模TF-DNA相互作用,基于对数线性模型。我们的方法使用序列特征来表示TF结合特异性,其中每个特征可以跨越多个位置。我们开发了我们的模型的数学公式,并设计了一种算法,用于学习其结构特征结合位点的数据。我们还开发了一种判别式基序发现器,其发现与背景组相比在靶序列组中富集的从头Festival。我们评估我们的方法合成的数据和广泛使用的TF染色质免疫沉淀(ChIP)数据集的Harbison等。然后,我们将我们的算法应用于高通量TF ChIP数据从小鼠和人类,揭示序列特征,存在于小鼠和人类的TF的结合特异性,并表明,FRENT解释TF结合显着优于PSSMs。我们的FMM学习和主题查找软件可在http://genie.weizmann.ac.il/上获得。转录因子(TF)蛋白与其DNA靶序列的结合是基因调控的基本物理相互作用。表征转录因子的结合特异性对于推断哪些基因受哪些转录因子调控是必不可少的。最近,开发了几种高通量的方法,测量富含TF靶基因组的序列。由于TF识别相对较短的序列,许多努力已经针对开发从这些序列中识别富集的序列(基序)的计算方法。然而,很少有人致力于改善图案的表现。实际上,可用的基序发现软件使用位置特异性评分矩阵(PSSM)模型,其假设不同基序位置之间的独立性。我们提出了一种替代的,更丰富的模型,称为特征基序模型(FMM),使各种序列特征的表示和捕获的依赖关系,存在于结合位点的位置之间。我们展示了如何在合成数据和真实的数据上,比PSSMs更好地解释TF结合数据。我们还提出了一种基序查找算法,该算法从未对齐的启动子序列中学习FMM基序,并展示了如何从人类TF c-Myc和CTCF的结合数据中学习从头FMM,揭示有关其结合特异性的有趣见解。
Transcription factor (TF) binding to its DNA target site is a fundamental regulatory interaction. The most common model used to represent TF binding specificities is a position specific scoring matrix (PSSM), which assumes independence between binding positions. However, in many cases, this simplifying assumption does not hold. Here, we present feature motif models (FMMs), a novel probabilistic method for modeling TF–DNA interactions, based on log-linear models. Our approach uses sequence features to represent TF binding specificities, where each feature may span multiple positions. We develop the mathematical formulation of our model and devise an algorithm for learning its structural features from binding site data. We also developed a discriminative motif finder, which discovers de novo FMMs that are enriched in target sets of sequences compared to background sets. We evaluate our approach on synthetic data and on the widely used TF chromatin immunoprecipitation (ChIP) dataset of Harbison et al. We then apply our algorithm to high-throughput TF ChIP data from mouse and human, reveal sequence features that are present in the binding specificities of mouse and human TFs, and show that FMMs explain TF binding significantly better than PSSMs. Our FMM learning and motif finder software are available at http://genie.weizmann.ac.il/. Transcription factor (TF) protein binding to its DNA target sequences is a fundamental physical interaction underlying gene regulation. Characterizing the binding specificities of TFs is essential for deducing which genes are regulated by which TFs. Recently, several high-throughput methods that measure sequences enriched for TF targets genomewide were developed. Since TFs recognize relatively short sequences, much effort has been directed at developing computational methods that identify enriched subsequences (motifs) from these sequences. However, little effort has been directed towards improving the representation of motifs. Practically, available motif finding software use the position specific scoring matrix (PSSM) model, which assumes independence between different motif positions. We present an alternative, richer model, called the feature motif model (FMM), that enables the representation of a variety of sequence features and captures dependencies that exist between binding site positions. We show how FMMs explain TF binding data better than PSSMs on both synthetic and real data. We also present a motif finder algorithm that learns FMM motifs from unaligned promoter sequences and show how de novo FMMs, learned from binding data of the human TFs c-Myc and CTCF, reveal intriguing insights about their binding specificities.
DOI: 10.1371/journal.pcbi.0030039
发表时间: 2007-03-23
影响因子: 4.3
作者:
Eden E;Lipson D;Yogev S;Yakhini Z
通讯作者: Yakhini Z
DOI: 10.1038/nbt1246
发表时间: 2006-11-01
影响因子: 46.9
作者:
Berger, Michael F.;Philippakis, Anthony A.;Bulyk, Martha L.
通讯作者: Bulyk, Martha L.
DOI: 10.1093/bioinformatics/bti402
发表时间: 2005-06-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Hong, PY;Liu, XS;Wong, WH
通讯作者: Wong, WH
DOI: 10.1093/nar/27.1.318
发表时间: 1999-01-01
影响因子: 14.9
作者:
Heinemeyer, T;Chen, X;Wingender, E
通讯作者: Wingender, E
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y