Extrapolating histone marks across developmental stages, tissues, and species: an enhancer prediction case study

Extrapolating histone marks across developmental stages, tissues, and species: an enhancer prediction case study
复制标题

DOI:
10.1186/s12864-015-1264-3
复制
发表时间:
2015-02-21
期刊:
影响因子:
4.4
通讯作者:
Capra, John A.
Capra, John A.
中科院分区:
生物学2区
文献类型:
--
作者:
Capra, John A.

文献摘要

被引文献

相似文献

背景:基因调控DNA的动态激活和失活产生驱动细胞系分化的表达变化。确定在发育转变过程中活跃的调控区域对于理解基因组如何指定复杂的发育程序以及这些过程如何在疾病中被破坏是必要的。基因调控动力学是由多种因素介导的,包括转录因子的结合、DNA和组蛋白的甲基化和乙酰化。TF结合、DNA和组蛋白修饰的全基因组图谱已经在许多细胞环境中生成;然而,考虑到动物发育的多样性和复杂性,这些数据只涵盖了细胞和发育背景的一小部分。因此,有必要使用现有的表观遗传学和功能基因组学数据来分析数千种尚未表征的环境。结果:为了研究组蛋白修饰数据在没有这些数据的细胞背景分析中的实用性,我评估了在不同发育阶段、组织和物种中收集的全基因组H3K27ac和H3K4me1数据能够在多大程度上预测实验验证的小鼠胚胎期11.5天(E11.5)活跃的心脏增强子。使用机器学习方法整合来自不同背景的数据,我发现E11.5心脏增强子通常可以从其他背景的数据中准确预测,并且我量化了每个数据源对预测的贡献。每个数据集的效用与发育时间和组织与目标环境的接近程度相关:来自发育后期和成人心脏组织的数据对于预测E11.5增强子最有帮助,而来自干细胞和早期发育阶段的标记信息较少。基于在非心脏组织和人类心脏中收集的数据的预测比随机预测要好,但比使用小鼠心脏数据的预测差。结论:这些算法基于相关但不同的细胞环境数据准确预测发育增强子的能力表明,将计算模型与从相关环境中采样的表观遗传数据相结合,可能足以实现许多感兴趣的细胞环境的功能表征。
Background: Dynamic activation and inactivation of gene regulatory DNA produce the expression changes that drive the differentiation of cellular lineages. Identifying regulatory regions active during developmental transitions is necessary to understand how the genome specifies complex developmental programs and how these processes are disrupted in disease. Gene regulatory dynamics are mediated by many factors, including the binding of transcription factors (TFs) and the methylation and acetylation of DNA and histones. Genome-wide maps of TF binding and DNA and histone modifications have been generated for many cellular contexts; however, given the diversity and complexity of animal development, these data cover only a small fraction of the cellular and developmental contexts of interest. Thus, there is a need for methods that use existing epigenetic and functional genomics data to analyze the thousands of contexts that remain uncharacterized.Results: To investigate the utility of histone modification data in the analysis of cellular contexts without such data, I evaluated how well genome-wide H3K27ac and H3K4me1 data collected in different developmental stages, tissues, and species were able to predict experimentally validated heart enhancers active at embryonic day 11.5 (E11.5) in mouse. Using a machine-learning approach to integrate the data from different contexts, I found that E11.5 heart enhancers can often be predicted accurately from data from other contexts, and I quantified the contribution of each data source to the predictions. The utility of each dataset correlated with nearness in developmental time and tissue to the target context: data from late developmental stages and adult heart tissues were most informative for predicting E11.5 enhancers, while marks from stem cells and early developmental stages were less informative. Predictions based on data collected in non-heart tissues and in human hearts were better than random, but worse than using data from mouse hearts.Conclusions: The ability of these algorithms to accurately predict developmental enhancers based on data from related, but distinct, cellular contexts suggests that combining computational models with epigenetic data sampled from relevant contexts may be sufficient to enable functional characterization of many cellular contexts of interest.