Stability selection for regression-based models of transcription factor-DNA binding specificity.

Stability selection for regression-based models of transcription factor-DNA binding specificity.
复制标题

DOI:
10.1093/bioinformatics/btt221
复制
发表时间:
2013-07-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Gordân R
Gordân R
中科院分区:
其他
文献类型:
--
作者:
Mordelet F;Horton J;Hartemink AJ;Engelhardt BE;Gordân R

文献摘要

参考文献

被引文献

相似文献

动机:转录因子(TF)的DNA结合特异性通常使用位置权重矩阵模型来表示,该模型隐含地假设TF结合位点上的单个碱基独立地贡献了结合亲和力,这一假设并不总是成立。由于这个原因,更复杂的结合特异性模型已经开发出来。然而,这些模型有它们自己的警告:它们通常有大量的参数,这使得它们很难学习和解释。结果:我们提出了新的基于回归的TF-DNA结合特异性模型,该模型使用自定义蛋白结合微阵列(PBM)实验的高分辨率体外数据进行训练。我们的PBMs是专门设计用于覆盖大量假定的DNA结合位点的目标TFs(酵母TFs Cbf1和Tye7,以及人类TFs c-Myc, Max和Mad2)在其原生基因组背景下。这些高通量的定量数据非常适合于训练复杂的模型,这些模型不仅考虑了单个碱基的独立贡献,而且考虑了结合位点内或附近不同位置的二核苷酸和三核苷酸的贡献。为了确保我们的模型保持可解释性,我们使用特征选择来识别少量序列特征,这些特征可以准确预测TF-DNA结合特异性。为了进一步说明我们回归模型的准确性,我们表明,即使在具有高度相似位置权重矩阵的平行TF的情况下,我们的新模型也可以区分单个因素的特异性。因此,我们的工作是朝着更好的基于序列的个体TF-DNA结合特异性模型迈出的重要一步。可用性:我们的代码可在http://genome.duke.edu/labs/gordan/ISMB2013上获得。本文中使用的PBM数据可在Gene Expression Omnibus中获得,登录号为GSE47026。联系人:raluca.gordan@duke.edu
Motivation: The DNA binding specificity of a transcription factor (TF) is typically represented using a position weight matrix model, which implicitly assumes that individual bases in a TF binding site contribute independently to the binding affinity, an assumption that does not always hold. For this reason, more complex models of binding specificity have been developed. However, these models have their own caveats: they typically have a large number of parameters, which makes them hard to learn and interpret. Results: We propose novel regression-based models of TF–DNA binding specificity, trained using high resolution in vitro data from custom protein-binding microarray (PBM) experiments. Our PBMs are specifically designed to cover a large number of putative DNA binding sites for the TFs of interest (yeast TFs Cbf1 and Tye7, and human TFs c-Myc, Max and Mad2) in their native genomic context. These high-throughput quantitative data are well suited for training complex models that take into account not only independent contributions from individual bases, but also contributions from di- and trinucleotides at various positions within or near the binding sites. To ensure that our models remain interpretable, we use feature selection to identify a small number of sequence features that accurately predict TF–DNA binding specificity. To further illustrate the accuracy of our regression models, we show that even in the case of paralogous TF with highly similar position weight matrices, our new models can distinguish the specificities of individual factors. Thus, our work represents an important step toward better sequence-based models of individual TF–DNA binding specificity. Availability: Our code is available at http://genome.duke.edu/labs/gordan/ISMB2013. The PBM data used in this article are available in the Gene Expression Omnibus under accession number GSE47026. Contact: raluca.gordan@duke.edu
DOI: 10.1126/science.1162327
发表时间: 2009-06-26
期刊: Science (New York, N.Y.)
影响因子: --
作者:
Badis G;Berger MF;Philippakis AA;Talukder S;Gehrke AR;Jaeger SA;Chan ET;Metzler G;Vedenko A;Chen X;Kuznetsov H;Wang CF;Coburn D;Newburger DE;Morris Q;Hughes TR;Bulyk ML
通讯作者: Bulyk ML
DOI: 10.1038/nprot.2008.195
发表时间: 2009
期刊: NATURE PROTOCOLS
影响因子: 14.8
作者:
Berger, Michael F.;Bulyk, Martha L.
通讯作者: Bulyk, Martha L.
DOI: 10.1371/journal.pone.0020059
发表时间: 2011
期刊: PloS one
影响因子: 3.7
作者:
Annala M;Laurila K;Lähdesmäki H;Nykter M
通讯作者: Nykter M
DOI: 10.1371/journal.pcbi.1000916
发表时间: 2010-09-01
影响因子: 4.3
作者:
Agius, Phaedra;Arvey, Aaron;Leslie, Christina
通讯作者: Leslie, Christina
DOI: 10.1038/nbt1246
发表时间: 2006-11-01
影响因子: 46.9
作者:
Berger, Michael F.;Philippakis, Anthony A.;Bulyk, Martha L.
通讯作者: Bulyk, Martha L.