Modelling the conditional regulatory activity of methylated and bivalent promoters.

Modelling the conditional regulatory activity of methylated and bivalent promoters.
复制标题

DOI:
10.1186/s13072-015-0013-9
复制
发表时间:
2015
影响因子:
3.9
通讯作者:
Crampin EJ
Crampin EJ
中科院分区:
生物学2区
文献类型:
--
作者:
Budden DM;Hurley DG;Crampin EJ

文献摘要

参考文献

被引文献

相似文献

基因表达的预测建模是通过整合高通量组学数据进行转录调控相互作用的计算机探索的强大框架。以前的方法的一个主要限制是它们无法处理当基因受制于不同的调节机制时出现的条件相互作用。尽管基于染色质免疫沉淀的组蛋白修饰数据通常被用作染色质可及性的代理,但这些变量与表达之间的关联通常取决于其他表观遗传标记(例如DNA甲基化或组蛋白变异)的存在。这些条件相互作用在以前的预测模型中处理得很差,降低了下游生物推理的可靠性。我们之前已经证明,整合转录因子和组蛋白修饰数据在一个单一的预测模型是无效的,因为它们的统计冗余。在本研究中,我们评估了四种量化基因水平DNA甲基化水平的建议方法,并证明在预测建模框架中包含这些数据也受到数据集成的关键限制。基于表观遗传数据中的统计冗余是由动态染色质背景下的条件调控相互作用引起的假设,我们构建了一个新的基因表达模型,该模型首次通过无监督识别潜在调控类来提高预测准确性。我们展示了DNA甲基化和H2A。Z组蛋白变异数据可以通过这种方式进行解释,以识别和探索沉默和二价启动子的特征,从而大大提高mRNA转录物丰度的全基因组预测和跨多个细胞系的下游生物学推断。先前的基因表达模型已经成功地应用于分子生物学中的几个重要问题,包括转录因子作用的发现,负责差异表达模式的调控元件的鉴定以及远距离物种转录组的比较分析。我们的分析支持了我们的假设,即表观遗传数据中的统计冗余部分是由于这些调节因子和基因表达水平之间的条件关系。该分析揭示了H3K4me3和H3K27me3在H2A存在下的异质作用。Z组蛋白变异(与癌症进展有关)以及这些特征在谱系承诺和癌变过程中如何变化。本文的在线版本(doi:10.1186/s13072-015-0013-9)包含补充材料,授权用户可以使用。
Predictive modelling of gene expression is a powerful framework for the in silico exploration of transcriptional regulatory interactions through the integration of high-throughput -omics data. A major limitation of previous approaches is their inability to handle conditional interactions that emerge when genes are subject to different regulatory mechanisms. Although chromatin immunoprecipitation-based histone modification data are often used as proxies for chromatin accessibility, the association between these variables and expression often depends upon the presence of other epigenetic markers (e.g. DNA methylation or histone variants). These conditional interactions are poorly handled by previous predictive models and reduce the reliability of downstream biological inference. We have previously demonstrated that integrating both transcription factor and histone modification data within a single predictive model is rendered ineffective by their statistical redundancy. In this study, we evaluate four proposed methods for quantifying gene-level DNA methylation levels and demonstrate that inclusion of these data in predictive modelling frameworks is also subject to this critical limitation in data integration. Based on the hypothesis that statistical redundancy in epigenetic data is caused by conditional regulatory interactions within a dynamic chromatin context, we construct a new gene expression model which is the first to improve prediction accuracy by unsupervised identification of latent regulatory classes. We show that DNA methylation and H2A.Z histone variant data can be interpreted in this way to identify and explore the signatures of silenced and bivalent promoters, substantially improving genome-wide predictions of mRNA transcript abundance and downstream biological inference across multiple cell lines. Previous models of gene expression have been applied successfully to several important problems in molecular biology, including the discovery of transcription factor roles, identification of regulatory elements responsible for differential expression patterns and comparative analysis of the transcriptome across distant species. Our analysis supports our hypothesis that statistical redundancy in epigenetic data is partially due to conditional relationships between these regulators and gene expression levels. This analysis provides insight into the heterogeneous roles of H3K4me3 and H3K27me3 in the presence of the H2A.Z histone variant (implicated in cancer progression) and how these signatures change during lineage commitment and carcinogenesis. The online version of this article (doi:10.1186/s13072-015-0013-9) contains supplementary material, which is available to authorized users.
DOI: 10.1016/j.stem.2011.12.017
发表时间: 2012-02-03
期刊: CELL STEM CELL
影响因子: 23.9
作者:
Brookes, Emily;de Santiago, Ines;Hebenstreit, Daniel;Morris, Kelly J.;Carroll, Tom;Xie, Sheila Q.;Stock, Julie K.;Heidemann, Martin;Eick, Dirk;Nozaki, Naohito;Kimura, Hiroshi;Ragoussis, Jiannis;Teichmann, Sarah A.;Pombo, Ana
通讯作者: Pombo, Ana
DOI: 10.1016/s0960-9822(00)00610-2
发表时间: 2000-07-27
期刊: CURRENT BIOLOGY
影响因子: 9.2
作者:
Paull, TT;Rogakou, EP;Bonner, WM
通讯作者: Bonner, WM
DOI: 10.1073/pnas.97.18.10101
发表时间: 2000-08-29
影响因子: 11.1
作者:
Alter, O;Brown, PO;Botstein, D
通讯作者: Botstein, D
DOI: 10.1093/nar/gkr752
发表时间: 2012-01
影响因子: 14.9
作者:
Cheng C;Gerstein M
通讯作者: Gerstein M
DOI: 10.1093/bioinformatics/bts529
发表时间: 2012-11-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
McLeay, Robert C.;Lesluyes, Tom;Bailey, Timothy L.
通讯作者: Bailey, Timothy L.