RFECS: a random-forest based algorithm for enhancer identification from chromatin state.

RFECS: a random-forest based algorithm for enhancer identification from chromatin state.
复制标题

DOI:
10.1371/journal.pcbi.1002968
复制
发表时间:
2013
影响因子:
4.3
通讯作者:
Ren B
Ren B
中科院分区:
生物学2区
文献类型:
--
作者:
Rajagopal N;Xie W;Li Y;Wagner U;Wang W;Stamatoyannopoulos J;Ernst J;Kellis M;Ren B

文献摘要

参考文献

被引文献

相似文献

转录增强子在基因表达调控中起着关键作用,但它们在真核基因组中的鉴定一直具有挑战性。最近,研究表明,哺乳动物基因组中的增强子与特征性组蛋白修饰模式相关,这已越来越多地用于增强子鉴定。然而,只有有限数量的细胞类型或染色质标记先前已被研究用于此目的,留下的问题没有回答是否存在一组最佳的组蛋白修饰的增强子预测在不同的细胞类型。在这里,我们通过探索两种不同的人类细胞类型(胚胎干细胞和肺成纤维细胞)中24种组蛋白修饰的全基因组概况来解决这个问题。我们开发了一种基于随机森林的算法RFECS(Random Forest based Enhancer identification from Chromatin States)来整合用于识别增强子的组蛋白修饰谱,并使用它来识别许多细胞类型中的增强子。我们表明,RFECS不仅导致更准确和精确的预测增强子比以前的方法,但也有助于确定信息最丰富和强大的一组三个染色质标记的增强子预测。增强子是基因组中可以激活基因表达的区域,而不管它们相对于基因的位置如何。识别这些元素对于理解不同细胞类型之间的调节差异至关重要。由于增强子缺乏特征性的序列特征,并且可能远离它们调节的基因,因此它们的识别并不简单。通过实验确定转录辅激活因子p300的全基因组结合位点是发现增强子的一种方法,但它只能识别增强子的一个子集。几年前,人们观察到p300的结合位点被独特的翻译后组蛋白修饰所标记。几个研究小组已经利用这一发现来预测全基因组增强子,基于它们与p300结合位点的组蛋白修饰谱的相似性。我们在这里报告了一种新的算法,用于此目的,并表明它具有更大的准确性比现有的方法。我们的算法的另一个独特的功能是能够自动推导出增强子预测所需的组蛋白修饰的信息量最大的集合。我们预计,随着已知组蛋白修饰数量的增加以及各种细胞类型和物种的表观基因组数据集的快速积累,这种方法将变得越来越有用。
Transcriptional enhancers play critical roles in regulation of gene expression, but their identification in the eukaryotic genome has been challenging. Recently, it was shown that enhancers in the mammalian genome are associated with characteristic histone modification patterns, which have been increasingly exploited for enhancer identification. However, only a limited number of cell types or chromatin marks have previously been investigated for this purpose, leaving the question unanswered whether there exists an optimal set of histone modifications for enhancer prediction in different cell types. Here, we address this issue by exploring genome-wide profiles of 24 histone modifications in two distinct human cell types, embryonic stem cells and lung fibroblasts. We developed a Random-Forest based algorithm, RFECS (Random Forest based Enhancer identification from Chromatin States) to integrate histone modification profiles for identification of enhancers, and used it to identify enhancers in a number of cell-types. We show that RFECS not only leads to more accurate and precise prediction of enhancers than previous methods, but also helps identify the most informative and robust set of three chromatin marks for enhancer prediction. Enhancers are regions in the genome that can activate the expression of a gene irrespective of their location with respect to the gene. Identifying these elements is critical in understanding regulatory differences between different cell-types. Since enhancers lack characteristic sequence features and can be far away from the gene they regulate, their identification is not trivial. Experimentally determining the genome-wide binding sites of transcriptional co-activator p300 is one way of finding enhancers but it can only identify a subset of enhancers. A few years ago, it was observed that the binding sites of p300 are marked by distinctive, post-translational histone modifications. Several groups have exploited this discovery to predict genome-wide enhancers based on their similarity to the histone modification profiles of p300 binding sites. We here report a novel algorithm for this purpose and show that it has much greater accuracy than existing methods. Another unique feature of our algorithm is the ability to automatically deduce the most informative set of histone modifications required for enhancer prediction. We expect that this method will become increasingly useful with the expanding number of known histone modifications and rapid accumulation of epigenomic datasets for various cell types and species.
DOI: 10.1038/nature09906
发表时间: 2011-05-05
期刊: NATURE
影响因子: 64.8
作者:
Ernst, Jason;Kheradpour, Pouya;Mikkelsen, Tarjei S.;Shoresh, Noam;Ward, Lucas D.;Epstein, Charles B.;Zhang, Xiaolan;Wang, Li;Issner, Robbyn;Coyne, Michael;Ku, Manching;Durham, Timothy;Kellis, Manolis;Bernstein, Bradley E.
通讯作者: Bernstein, Bradley E.
DOI: 10.1371/journal.pbio.1001046
发表时间: 2011-04
期刊: PLoS biology
影响因子: 9.8
作者:
ENCODE Project Consortium
通讯作者: ENCODE Project Consortium
DOI: 10.1038/nature07829
发表时间: 2009-05-07
期刊: NATURE
影响因子: 64.8
作者:
Heintzman, Nathaniel D.;Hon, Gary C.;Hawkins, R. David;Kheradpour, Pouya;Stark, Alexander;Harp, Lindsey F.;Ye, Zhen;Lee, Leonard K.;Stuart, Rhona K.;Ching, Christina W.;Ching, Keith A.;Antosiewicz-Bourget, Jessica E.;Liu, Hui;Zhang, Xinmin;Green, Roland D.;Lobanenkov, Victor V.;Stewart, Ron;Thomson, James A.;Crawford, Gregory E.;Kellis, Manolis;Ren, Bing
通讯作者: Ren, Bing
DOI: 10.1073/pnas.1016959108
发表时间: 2011-04-05
影响因子: 11.1
作者:
He, Aibin;Kong, Sek Won;Pu, William T.
通讯作者: Pu, William T.
DOI: 10.1073/pnas.1017214108
发表时间: 2011-03-29
影响因子: 11.1
作者:
Jin, Fulai;Li, Yan;Natarajan, Rama
通讯作者: Natarajan, Rama