Annotation of gene promoters by integrative data-mining of ChIP-seq Pol-II enrichment data.

Annotation of gene promoters by integrative data-mining of ChIP-seq Pol-II enrichment data.
复制标题

DOI:
10.1186/1471-2105-11-s1-s65
复制
发表时间:
2010-01-18
期刊:
影响因子:
3
通讯作者:
Davuluri RV
Davuluri RV
中科院分区:
生物学4区
文献类型:
--
作者:
Gupta R;Wikramasinghe P;Bhattacharyya A;Perez FA;Pal S;Davuluri RV

文献摘要

被引文献

相似文献

在哺乳动物基因组中,使用替代基因启动子驱动广泛的细胞型、组织型或发育型基因调控是一种常见现象。染色质免疫沉淀方法与DNA微阵列(ChIP-chip)或大规模平行测序(ChIP-seq)相结合,可以利用针对Pol-II的抗体在不同细胞条件下对活性启动子进行全基因组鉴定。然而,由于ChIP中使用的抗体的非特异性,这些方法不仅在基因启动子附近富集,而且在基因和其他基因组区域内富集。此外,这些方法的使用受到其高成本和对细胞类型和环境的强烈依赖的限制。我们训练并测试了不同的最先进的集合和元分类方法来识别Pol-II富集启动子和Pol-II富集非启动子序列,每个序列的长度为500 bp。分类模型在一个基准数据集上进行训练和测试,使用一组39个不同的特征变量,这些特征变量基于染色质修饰特征和各种DNA序列特征。将表现最好的模型应用于7个已发表的ChIP-seq Pol-II数据集,以提供小鼠基因启动子的全基因组注释。我们提出了一种基于监督学习方法的新算法,以区分启动子相关的Pol-II富集与ChIP-chip/seq图谱中基因组其他位置的富集。我们从五个组织(脑、肾、肝、肺和脾)的RNA Pol-II ChIP-seq数据中积累了11,773个启动子序列和46,167个非启动子序列的数据集,每个序列长度为500 bp。我们评估了建立最佳预测器的分类模型,发现基于Bagging和随机森林的方法具有最佳精度。我们在七个不同的ChIP-seq数据集上实现了该算法,为小鼠基因组中的蛋白质编码和非编码基因提供了一套全面的启动子注释。由此产生的注释包含13,413(4,747)个具有单个启动子的蛋白质编码(非编码)基因和9,929(1,858)个具有两个或多个可选启动子的蛋白质编码(非编码)基因,以及大量未指定的新启动子。我们的新算法可以成功地从Pol-II结合区域的全基因组谱中预测启动子。此外,我们的算法明显优于现有的启动子预测方法,可以应用于Pol-II启动子的全基因组预测。
Use of alternative gene promoters that drive widespread cell-type, tissue-type or developmental gene regulation in mammalian genomes is a common phenomenon. Chromatin immunoprecipitation methods coupled with DNA microarray (ChIP-chip) or massive parallel sequencing (ChIP-seq) are enabling genome-wide identification of active promoters in different cellular conditions using antibodies against Pol-II. However, these methods produce enrichment not only near the gene promoters but also inside the genes and other genomic regions due to the non-specificity of the antibodies used in ChIP. Further, the use of these methods is limited by their high cost and strong dependence on cellular type and context. We trained and tested different state-of-art ensemble and meta classification methods for identification of Pol-II enriched promoter and Pol-II enriched non-promoter sequences, each of length 500 bp. The classification models were trained and tested on a bench-mark dataset, using a set of 39 different feature variables that are based on chromatin modification signatures and various DNA sequence features. The best performing model was applied on seven published ChIP-seq Pol-II datasets to provide genome wide annotation of mouse gene promoters. We present a novel algorithm based on supervised learning methods to discriminate promoter associated Pol-II enrichment from enrichment elsewhere in the genome in ChIP-chip/seq profiles. We accumulated a dataset of 11,773 promoter and 46,167 non-promoter sequences, each of length 500 bp, generated from RNA Pol-II ChIP-seq data of five tissues (Brain, Kidney, Liver, Lung and Spleen). We evaluated the classification models in building the best predictor and found that Bagging and Random Forest based approaches give the best accuracy. We implemented the algorithm on seven different published ChIP-seq datasets to provide a comprehensive set of promoter annotations for both protein-coding and non-coding genes in the mouse genome. The resulting annotations contain 13,413 (4,747) protein-coding (non-coding) genes with single promoters and 9,929 (1,858) protein-coding (non-coding) genes with two or more alternative promoters, and a significant number of unassigned novel promoters. Our new algorithm can successfully predict the promoters from the genome wide profile of Pol-II bound regions. In addition, our algorithm performs significantly better than existing promoter prediction methods and can be applied for genome-wide predictions of Pol-II promoters.