DEEP MOTIF DASHBOARD: VISUALIZING AND UNDERSTANDING GENOMIC SEQUENCES USING DEEP NEURAL NETWORKS.

DEEP MOTIF DASHBOARD: VISUALIZING AND UNDERSTANDING GENOMIC SEQUENCES USING DEEP NEURAL NETWORKS.
复制标题

深图式仪表板:使用深神经网络可视化和理解基因组序列。

DOI:
10.1142/9789813207813_0025
复制
发表时间:
2017
影响因子:
--
通讯作者:
Qi Y
Qi Y
中科院分区:
其他
文献类型:
--
作者:
Lanchantin J;Singh R;Wang B;Qi Y

文献摘要

参考文献

被引文献

相似文献

深度神经网络(DNN)模型最近在转录因子结合(TFBS)位点分类任务中获得了最先进的预测精度。然而,目前尚不清楚这些方法如何识别有意义的DNA序列信号,以及为什么tf与特定位置结合。在本文中,我们提出了一个名为Deep Motif Dashboard (DeMo Dashboard)的工具包,它提供了一套可视化策略来从深度神经网络模型中提取用于TFBS分类的Motif或序列模式。我们演示了如何可视化和理解三个重要的深度神经网络模型:卷积、循环和卷积-循环网络。我们的第一种可视化方法是找到一个测试序列的显著性图,该图使用一阶导数来描述每个核苷酸在最终预测中的重要性。其次,考虑到循环模型以时间方式进行预测(从TFBS序列的一端到另一端),我们引入了时间输出分数,表示模型随时间的预测分数。最后,针对特定类别的可视化策略,通过随机梯度优化找到给定TFBS正类别的最优输入序列。实验结果表明,在三种结构中,卷积-循环结构的性能最好。可视化技术表明,CNN-RNN通过建模基序以及它们之间的依赖关系来进行预测。
Deep neural network (DNN) models have recently obtained state-of-the-art prediction accuracy for the transcription factor binding (TFBS) site classification task. However, it remains unclear how these approaches identify meaningful DNA sequence signals and give insights as to why TFs bind to certain locations. In this paper, we propose a toolkit called the Deep Motif Dashboard (DeMo Dashboard) which provides a suite of visualization strategies to extract motifs, or sequence patterns from deep neural network models for TFBS classification. We demonstrate how to visualize and understand three important DNN models: convolutional, recurrent, and convolutional-recurrent networks. Our first visualization method is finding a test sequence’s saliency map which uses first-order derivatives to describe the importance of each nucleotide in making the final prediction. Second, considering recurrent models make predictions in a temporal manner (from one end of a TFBS sequence to the other), we introduce temporal output scores, indicating the prediction score of a model over time for a sequential input. Lastly, a class-specific visualization strategy finds the optimal input sequence for a given TFBS positive class via stochastic gradient optimization. Our experimental results indicate that a convolutional-recurrent architecture performs the best among the three architectures. The visualization techniques indicate that CNN-RNN makes predictions by modeling both motifs as well as dependencies among them.
DOI: 10.1093/bioinformatics/16.1.16
发表时间: 2000-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Stormo, GD
通讯作者: Stormo, GD
DOI: 10.1093/bioinformatics/btw427
发表时间: 2016-09-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Singh, Ritambhara;Lanchantin, Jack;Qi, Yanjun
通讯作者: Qi, Yanjun
量化基序之间的相似性。
DOI: 10.1186/gb-2007-8-2-r24
发表时间: 2007
期刊: Genome biology
影响因子: 12.3
作者:
Gupta S;Stamatoyannopoulos JA;Bailey TL;Noble WS
通讯作者: Noble WS
DOI: 10.1145/3065386
发表时间: 2017-06-01
影响因子: 22.7
作者:
Krizhevsky, Alex;Sutskever, Ilya;Hinton, Geoffrey E.
通讯作者: Hinton, Geoffrey E.
DOI: 10.1093/bioinformatics/btr189
发表时间: 2011-06-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Machanick P;Bailey TL
通讯作者: Bailey TL