Predicting enhancers with deep convolutional neural networks.

Predicting enhancers with deep convolutional neural networks.
复制标题

使用深度卷积神经网络预测增强器

DOI:
10.1186/s12859-017-1878-3
复制
发表时间:
2017-12-01
期刊:
影响因子:
3
通讯作者:
Jiang R
Jiang R
中科院分区:
生物学4区
文献类型:
--
作者:
Min X;Zeng W;Chen S;Chen N;Chen T;Jiang R

文献摘要

参考文献

被引文献

相似文献

背景近年来,随着深度测序技术的快速发展,增强子在FANTOM和ENCODE等项目中得到系统鉴定,在一系列人类细胞系中形成了全基因组景观。然而,实验方法仍然是昂贵的和耗时的大规模识别增强子在各种组织在不同的疾病状态下,使计算识别的增强子important.ResultsTo促进增强子的识别,我们提出了一个计算框架,命名为DeepEnhancer,区分增强子从背景基因组序列。我们的方法纯粹依赖于DNA序列,通过使用深度卷积神经网络(CNN)以端到端的方式预测增强子。我们在许可增强子上训练我们的深度学习模型,然后采用迁移学习策略来微调特定于细胞系的增强子上的模型。结果表明,我们的方法在针对随机序列的增强子分类中的有效性和效率,表现出深度学习优于传统的基于序列的分类器的优势。然后,我们构建了各种不同的架构的神经网络,并显示在我们的方法中的最大池和批量归一化等技术的有用性。为了获得我们方法的可解释性,我们进一步将卷积核可视化为序列标识,并成功识别JASPAR数据库中的相似基序。ConclusionsDeepEnhancer通过高度准确的深度学习模型,仅使用DNA序列即可识别新型增强子。所提出的计算框架也可以应用于类似的问题,从而促进机器学习方法在生命科学中的使用。
BackgroundWith the rapid development of deep sequencing techniques in the recent years, enhancers have been systematically identified in such projects as FANTOM and ENCODE, forming genome-wide landscapes in a series of human cell lines. Nevertheless, experimental approaches are still costly and time consuming for large scale identification of enhancers across a variety of tissues under different disease status, making computational identification of enhancers indispensable.ResultsTo facilitate the identification of enhancers, we propose a computational framework, named DeepEnhancer, to distinguish enhancers from background genomic sequences. Our method purely relies on DNA sequences to predict enhancers in an end-to-end manner by using a deep convolutional neural network (CNN). We train our deep learning model on permissive enhancers and then adopt a transfer learning strategy to fine-tune the model on enhancers specific to a cell line. Results demonstrate the effectiveness and efficiency of our method in the classification of enhancers against random sequences, exhibiting advantages of deep learning over traditional sequence-based classifiers. We then construct a variety of neural networks with different architectures and show the usefulness of such techniques as max-pooling and batch normalization in our method. To gain the interpretability of our approach, we further visualize convolutional kernels as sequence logos and successfully identify similar motifs in the JASPAR database.ConclusionsDeepEnhancer enables the identification of novel enhancers using only DNA sequences via a highly accurate deep learning model. The proposed computational framework can also be applied to similar problems, thereby prompting the use of machine learning methods in life sciences.
DOI: 10.1038/ng.2892
发表时间: 2014-03
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Kircher, Martin;Witten, Daniela M.;Jain, Preti;O'Roak, Brian J.;Cooper, Gregory M.;Shendure, Jay
通讯作者: Shendure, Jay
模因套件:用于发现和搜索的工具。
DOI: 10.1093/nar/gkp335
发表时间: 2009-07
影响因子: 14.9
作者:
Bailey TL;Boden M;Buske FA;Frith M;Grant CE;Clementi L;Ren J;Li WW;Noble WS
通讯作者: Noble WS
DOI: 10.1093/bioinformatics/btx234
发表时间: 2017-07-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Min X;Zeng W;Chen N;Chen T;Jiang R
通讯作者: Jiang R
DOI: 10.1093/nar/gkv1176
发表时间: 2016-01-04
影响因子: 14.9
作者:
Mathelier A;Fornes O;Arenillas DJ;Chen CY;Denay G;Lee J;Shi W;Shyr C;Tan G;Worsley-Hunt R;Zhang AW;Parcy F;Lenhard B;Sandelin A;Wasserman WW
通讯作者: Wasserman WW
DOI: 10.1101/gr.5704207
发表时间: 2007-06-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Koch, Christoph M.;Andrews, Robert M.;Dunham, Ian
通讯作者: Dunham, Ian