Opening up the blackbox: an interpretable deep neural network-based classifier for cell-type specific enhancer predictions.

Opening up the blackbox: an interpretable deep neural network-based classifier for cell-type specific enhancer predictions.
复制标题

DOI:
10.1186/s12918-016-0302-3
复制
发表时间:
2016-08-01
影响因子:
--
通讯作者:
Chaterji S
Chaterji S
中科院分区:
生物2区
文献类型:
--
作者:
Kim SG;Theera-Ampornpunt N;Fang CH;Harwani M;Grama A;Chaterji S

文献摘要

被引文献

相似文献

基因表达是由专门的顺式调控模块(CRM),其中最突出的是所谓的增强子介导。早期的实验表明,位于远离基因启动子的增强子通常负责介导基因转录。了解它们的特性、调节活性和基因组靶点对于从细胞稳态到分化的细胞事件的功能理解至关重要。最近对表观基因组标记的全基因组研究表明,增强子元件可以富集某些表观基因组标记,例如组蛋白修饰的组合模式。我们在这篇论文中的努力是出于这些最新进展的表观基因组分析方法,它揭示了增强子相关的染色质功能,在不同的细胞类型和生物体。具体来说,在本文中,我们使用最新的深度学习方法,并开发了一种基于深度神经网络(DNN)的架构,称为EP-DNN,以预测人类基因组中增强子的存在和类型。它使用功能位点的峰处以及其邻近区域中的组蛋白修饰的表达水平作为特征。我们将EP-DNN应用于四种不同的细胞类型:H1,IMR 90,HepG 2和HeLa S3。我们使用p300结合位点作为增强子,TSS和随机非DHS位点作为非增强子来训练EP-DNN。我们执行EP-DNN预测以量化预测中不同置信度的验证率,并与两种最先进的增强子预测计算模型DEEP-ENCODE和RFECS进行比较。我们发现,EP-DNN具有上级的准确性,并需要更少的时间来进行预测。接下来,我们开发了通过计算分类任务中每个输入特征的重要性来使EP-DNN可解释的方法。该分析表明,重要的组蛋白修饰对于不同的细胞类型是不同的,具有一些重叠,例如,H3 K27 ac在H1细胞中起重要作用,但在HeLa S3细胞中不那么重要,而H3 K4 me 1在所有四种细胞类型中相对重要。最后,我们使用特征重要性分析来减少训练DNN所需的输入特征数量,从而减少训练时间,这通常是DNN使用中的计算瓶颈。在本文中,我们开发了EP-DNN,它具有高预测准确性,对于我们研究的所有四种细胞系的增强子预测的操作区域,验证率超过90%,优于DEEP-ENCODE和RFECS。然后,我们开发了一种方法来分析经过训练的DNN,并确定哪些组蛋白修饰是重要的,以及增强子位点近端或远端的哪些特征是重要的。
Gene expression is mediated by specialized cis-regulatory modules (CRMs), the most prominent of which are called enhancers. Early experiments indicated that enhancers located far from the gene promoters are often responsible for mediating gene transcription. Knowing their properties, regulatory activity, and genomic targets is crucial to the functional understanding of cellular events, ranging from cellular homeostasis to differentiation. Recent genome-wide investigation of epigenomic marks has indicated that enhancer elements could be enriched for certain epigenomic marks, such as, combinatorial patterns of histone modifications. Our efforts in this paper are motivated by these recent advances in epigenomic profiling methods, which have uncovered enhancer-associated chromatin features in different cell types and organisms. Specifically, in this paper, we use recent state-of-the-art Deep Learning methods and develop a deep neural network (DNN)-based architecture, called EP-DNN, to predict the presence and types of enhancers in the human genome. It uses as features, the expression levels of the histone modifications at the peaks of the functional sites as well as in its adjacent regions. We apply EP-DNN to four different cell types: H1, IMR90, HepG2, and HeLa S3. We train EP-DNN using p300 binding sites as enhancers, and TSS and random non-DHS sites as non-enhancers. We perform EP-DNN predictions to quantify the validation rate for different levels of confidence in the predictions and also perform comparisons against two state-of-the-art computational models for enhancer predictions, DEEP-ENCODE and RFECS. We find that EP-DNN has superior accuracy and takes less time to make predictions. Next, we develop methods to make EP-DNN interpretable by computing the importance of each input feature in the classification task. This analysis indicates that the important histone modifications were distinct for different cell types, with some overlaps, e.g., H3K27ac was important in cell type H1 but less so in HeLa S3, while H3K4me1 was relatively important in all four cell types. We finally use the feature importance analysis to reduce the number of input features needed to train the DNN, thus reducing training time, which is often the computational bottleneck in the use of a DNN. In this paper, we developed EP-DNN, which has high accuracy of prediction, with validation rates above 90 % for the operational region of enhancer prediction for all four cell lines that we studied, outperforming DEEP-ENCODE and RFECS. Then, we developed a method to analyze a trained DNN and determine which histone modifications are important, and within that, which features proximal or distal to the enhancer site, are important.