Chromatin accessibility prediction via convolutional long short-term memory networks with k-mer embedding.

Chromatin accessibility prediction via convolutional long short-term memory networks with k-mer embedding.
复制标题

DOI:
10.1093/bioinformatics/btx234
复制
发表时间:
2017-07-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Jiang R
Jiang R
中科院分区:
其他
文献类型:
--
作者:
Min X;Zeng W;Chen N;Chen T;Jiang R

文献摘要

参考文献

被引文献

相似文献

用于测量染色质可及性的实验技术是昂贵且耗时的,吸引了计算方法的发展以从DNA序列预测开放的染色质区域。沿着这个方向,现有的方法分为两类:一类基于手工制作的k-mer特征,另一类基于卷积神经网络。尽管到目前为止,这两个类别在特定应用中表现出良好的性能,但仍然缺乏一个全面的框架来整合有用的k-mer共现信息与深度学习的最新进展。我们通过使用具有k-mer嵌入的卷积长短期记忆(LSTM)网络来解决染色质可及性预测问题来填补这一空白。我们首先将DNA序列分成k-mer,并通过使用无监督表示学习方法基于k-mer的共生矩阵预训练k-mer嵌入向量。然后,我们构建了一个有监督的深度学习架构,该架构由一个嵌入层、三个卷积层和一个用于特征学习和分类的双向LSTM(BLSTM)层组成。我们证明了我们的方法从可变长度序列中获得高质量的固定长度特征,并且始终优于基线方法。我们发现,k-mer嵌入可以有效地提高模型的性能,通过探索不同的嵌入策略。我们还证明了卷积和BLSTM层的效率,通过比较两种变化的网络架构。我们通过敏感性分析证实了我们的模型对超参数的鲁棒性。我们希望我们的方法最终能够加强我们对在基因组研究中使用深度学习的理解,并为染色质可及性机制的研究提供帮助。源代码可以从https://github.com/minxueric/ismb2017_lstm下载。 补充材料可在生物信息学在线。
Experimental techniques for measuring chromatin accessibility are expensive and time consuming, appealing for the development of computational approaches to predict open chromatin regions from DNA sequences. Along this direction, existing methods fall into two classes: one based on handcrafted k-mer features and the other based on convolutional neural networks. Although both categories have shown good performance in specific applications thus far, there still lacks a comprehensive framework to integrate useful k-mer co-occurrence information with recent advances in deep learning. We fill this gap by addressing the problem of chromatin accessibility prediction with a convolutional Long Short-Term Memory (LSTM) network with k-mer embedding. We first split DNA sequences into k-mers and pre-train k-mer embedding vectors based on the co-occurrence matrix of k-mers by using an unsupervised representation learning approach. We then construct a supervised deep learning architecture comprised of an embedding layer, three convolutional layers and a Bidirectional LSTM (BLSTM) layer for feature learning and classification. We demonstrate that our method gains high-quality fixed-length features from variable-length sequences and consistently outperforms baseline methods. We show that k-mer embedding can effectively enhance model performance by exploring different embedding strategies. We also prove the efficacy of both the convolution and the BLSTM layers by comparing two variations of the network architecture. We confirm the robustness of our model to hyper-parameters by performing sensitivity analysis. We hope our method can eventually reinforce our understanding of employing deep learning in genomic studies and shed light on research regarding mechanisms of chromatin accessibility. The source code can be downloaded from https://github.com/minxueric/ismb2017_lstm. Supplementary materials are available at Bioinformatics online.
DOI: 10.1038/ng.759
发表时间: 2011-03
期刊: Nature genetics
影响因子: 30.8
作者:
John S;Sabo PJ;Thurman RE;Sung MH;Biddie SC;Johnson TA;Hager GL;Stamatoyannopoulos JA
通讯作者: Stamatoyannopoulos JA
DOI: 10.1016/s0893-6080(03)00138-2
发表时间: 2003-12-01
期刊: NEURAL NETWORKS
影响因子: 7.8
作者:
Wilson, DR;Martinez, TR
通讯作者: Martinez, TR
DOI: 10.1101/gr.121905.111
发表时间: 2011-12-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Lee, Dongwon;Karchin, Rachel;Beer, Michael A.
通讯作者: Beer, Michael A.
DOI: 10.1109/72.279181
发表时间: 1994-03-01
影响因子: --
作者:
BENGIO, Y;SIMARD, P;FRASCONI, P
通讯作者: FRASCONI, P
DOI: 10.1101/gad.1615707
发表时间: 2007-11-01
影响因子: 10.5
作者:
Niwa, Hitoshi
通讯作者: Niwa, Hitoshi