A Pretraining-Retraining Strategy of Deep Learning Improves Cell-Specific Enhancer Predictions

A Pretraining-Retraining Strategy of Deep Learning Improves Cell-Specific Enhancer Predictions
复制标题

深度学习的预训练-再训练策略改善了细胞特异性增强子预测

DOI:
10.3389/fgene.2019.01305
复制
发表时间:
2020-01
影响因子:
3.7
通讯作者:
Xuehai Hu
Xuehai Hu
中科院分区:
生物学3区
文献类型:
--
作者:
Xiaohui Niu;Kun Yang;Ge Zhang;Zhiquan Yang;Xuehai Hu

文献摘要

参考文献

相似文献

解析顺式调控元件(CRE)的编码是当今生物学的核心问题之一。增强子是远端顺式调控元件,在基因转录调控中发挥着重要作用。尽管识别全基因组范围内增强子的位置[区分性增强子预测(DEP)]很有必要,但预测增强子会在哪些特定细胞或组织类型中被激活并发挥功能[组织特异性增强子预测(TSEP)]更为重要。尽管现有的深度学习模型在DEP方面取得了巨大成功,但它们不能直接应用于TSEP,因为特定的细胞或组织类型仅有数量有限的增强子样本可用于训练。在此,我们首先采用一种已报道的深度学习架构,然后通过将整个训练过程分解为两个连续阶段,开发出一种名为“预训练 - 再训练策略”(PRS)的新型训练策略用于TSEP:预训练阶段旨在利用全部增强子数据进行训练以执行DEP,随后的再训练策略则基于预训练模型,利用组织特异性增强子样本进行训练以实现TSEP。结果表明,当通过五折交叉验证在更大规模的FANTOM5增强子数据集上进行测试时,PRS在DEP方面有效,其曲线下面积(AUC)为0.922,几何均值(GM)为0.696。有趣的是,基于训练好的预训练模型,一个新发现是,在测试23种特定组织或细胞系时,仅需额外20个训练轮次即可完成再训练过程。对于TSEP任务,PRS实现了0.806的平均GM,显著高于现有CRE预测主流方法gkm - SVM的0.528。值得注意的是,PRS进一步被证明优于另外两种最先进的方法:DEEP和BiRen。总之,PRS借鉴了迁移学习领域的有用理念,是一种用于TSEP的可靠方法。
Deciphering the code of cis-regulatory element (CRE) is one of the core issues of today ’ s.biology. Enhancers are distal CREs and play signi fi cant roles in gene transcriptional.regulation. Although identi fi cations of enhancer locations across the whole genome.[discriminative enhancer predictions (DEP)] is necessary, it is more important to predict.in which speci fi c cell or tissue types, they will be activated and functional [tissue-speci fi c.enhancer predictions (TSEP)]. Although existing deep learning models achieved great.successes in DEP, they cannot be directly employed in TSEP because a speci fi c cell or.tissue type only has a limited number of available enhancer samples for training. Here, we.fi rst adopted a reported deep learning architecture and then developed a novel training.strategy named “ pretraining-retraining strategy ” (PRS) for TSEP by decomposing the.whole training process into two successive stages: a pretraining stage is designed to train.with the whole enhancer data for performing DEP, and a retraining strategy is then.designed to train with tissue-speci fi c enhancer samples based on the trained pretraining.model for making TSEP. As a result, PRS is found to be valid for DEP with an AUC of 0.922.and a GM (geometric mean) of 0.696, when testing on a larger-scale FANTOM5 enhancer.dataset via a fi ve-fold cross-validation. Interestingly, based on the trained pretraining.model, a new fi nding is that only additional twenty epochs are needed to complete the.retraining process on testing 23 speci fi c tissues or cell lines. For TSEP tasks, PRS.achieved a mean GM of 0.806 which is signi fi cantly higher than 0.528 of gkm-SVM, an.existing mainstream method for CRE predictions. Notably, PRS is further proven superior.to other two state-of-the-art methods: DEEP and BiRen. In summary, PRS has employed.useful ideas from the domain of transfer learning and is a reliable method for TSEPs.
DOI: --
发表时间: --
影响因子: 2.9
作者:
Ying Huang;Beifang Niu;Ying Gao;L. Fu;Weizhong Li
通讯作者: Ying Huang;Beifang Niu;Ying Gao;L. Fu;Weizhong Li
DOI: 10.1038/nature07829
发表时间: 2009-05-07
期刊: NATURE
影响因子: 64.8
作者:
Heintzman, Nathaniel D.;Hon, Gary C.;Hawkins, R. David;Kheradpour, Pouya;Stark, Alexander;Harp, Lindsey F.;Ye, Zhen;Lee, Leonard K.;Stuart, Rhona K.;Ching, Christina W.;Ching, Keith A.;Antosiewicz-Bourget, Jessica E.;Liu, Hui;Zhang, Xinmin;Green, Roland D.;Lobanenkov, Victor V.;Stewart, Ron;Thomson, James A.;Crawford, Gregory E.;Kellis, Manolis;Ren, Bing
通讯作者: Ren, Bing
DOI: 10.1101/gr.173518.114
发表时间: 2014-10
期刊: Genome research
影响因子: 7
作者:
Kwasnieski JC;Fiore C;Chaudhari HG;Cohen BA
通讯作者: Cohen BA
DOI: 10.4018/978-1-7998-1192-3.ch008
发表时间: 2020
期刊: Advances in Systems Analysis, Software Engineering, and High Performance Computing
影响因子: --
作者:
Menaga D.;R. S.
通讯作者: Menaga D.;R. S.
DOI: 10.4018/978-1-5225-9096-5.ch007
发表时间: 2021-07
期刊: Smart Computational Intelligence in Biomedical and Health Informatics
影响因子: --
作者:
A. Sinha;S. Gupta;Anurag Tiwari;Amrita Chaturvedi
通讯作者: A. Sinha;S. Gupta;Anurag Tiwari;Amrita Chaturvedi