Attend and Predict: Understanding Gene Regulation by Selective Attention on Chromatin

Attend and Predict: Understanding Gene Regulation by Selective Attention on Chromatin
复制标题

DOI:
10.1101/329334
复制
发表时间:
2017-08
期刊:
bioRxiv
影响因子:
--
通讯作者:
Ritambhara Singh;Jack Lanchantin;Arshdeep Sekhon;Yanjun Qi
Ritambhara Singh;Jack Lanchantin;Arshdeep Sekhon;Yanjun Qi
中科院分区:
其他
文献类型:
--
作者:
Ritambhara Singh;Jack Lanchantin;Arshdeep Sekhon;Yanjun Qi

文献摘要

相似文献

过去十年见证了基因组技术的革命,使得染色质标记的全基因组分析成为可能。最近的文献试图通过大规模染色质测量预测基因表达来了解基因调控。此类学习任务存在两个基本挑战:(1)全基因组染色质信号具有空间结构、高维度和高度模块化; (2) 核心目标是了解相关因素是什么以及它们如何协同作用。以前的研究要么未能对输入信号之间的复杂依赖性进行建模,要么依赖单独的特征分析来解释决策。本文提出了一种基于注意力的深度学习方法 AttentiveChrome,该方法使用统一的架构来建模和解释染色质因子之间的依赖性,以控制基因调控。 AttentiveChrome 使用多个长短期记忆 (LSTM) 模块的层次结构对输入信号进行编码,并模拟各种染色质标记如何自动协作。 AttentiveChrome 与目标预测联合训练两个级别的注意力,使其能够差异化地关注相关标记并定位每个标记的重要位置。我们评估了人类 56 种不同细胞类型(任务)的模型。所提出的架构不仅更准确,而且其注意力分数比最先进的特征可视化方法(例如显着图)提供了更好的解释。1
The past decade has seen a revolution in genomic technologies that enabled a flood of genome-wide profiling of chromatin marks. Recent literature tried to understand gene regulation by predicting gene expression from large-scale chromatin measurements. Two fundamental challenges exist for such learning tasks: (1) genome-wide chromatin signals are spatially structured, high-dimensional and highly modular; and (2) the core aim is to understand what the relevant factors are and how they work together. Previous studies either failed to model complex dependencies among input signals or relied on separate feature analysis to explain the decisions. This paper presents an attention-based deep learning approach, AttentiveChrome, that uses a unified architecture to model and to interpret dependencies among chromatin factors for controlling gene regulation. AttentiveChrome uses a hierarchy of multiple Long Short-Term Memory (LSTM) modules to encode the input signals and to model how various chromatin marks cooperate automatically. AttentiveChrome trains two levels of attention jointly with the target prediction, enabling it to attend differentially to relevant marks and to locate important positions per mark. We evaluate the model across 56 different cell types (tasks) in humans. Not only is the proposed architecture more accurate, but its attention scores provide a better interpretation than state-of-the-art feature visualization methods such as saliency maps.1