Exploiting deep transfer learning for the prediction of functional non-coding variants using genomic sequence

Exploiting deep transfer learning for the prediction of functional non-coding variants using genomic sequence
复制标题

DOI:
10.1093/bioinformatics/btac214
复制
发表时间:
2022-05-13
期刊:
影响因子:
5.8
通讯作者:
Zhao, Fengdi
Zhao, Fengdi
中科院分区:
生物学3区
文献类型:
--
作者:
Chen, Li;Wang, Ye;Zhao, Fengdi

文献摘要

被引文献

相似文献

动机:虽然全基因组关联研究已经确定了数万种与复杂性状相关的变异,其中大多数属于非编码区,但它们可能不是因果关系。高通量功能测定的发展导致实验验证的非编码功能变体的发现。然而,由于技术难度和经济成本,这些经过验证的变体很少见。验证变体的样本量较小,因此开发监督机器学习模型来实现非编码因果变体的全基因组预测的可靠性较低。结果:我们将利用基于卷积神经网络的深度迁移学习模型来改善功能性非编码变体(NCV)的预测。为了解决小样本量的挑战,迁移学习模型利用大规模通用功能NCV来改善对低级特征的学习,并利用特定于上下文的功能NCV来学习高级特征以实现特定于上下文的预测任务。通过在三个MPRA数据集和16个GWAS数据集上评估深度迁移学习模型,我们证明了所提出的模型在没有预训练或再训练的情况下优于深度学习模型。此外,深度迁移学习模型在MPRA和GWAS数据集中的表现优于18种现有的计算方法。
Motivation: Though genome-wide association studies have identified tens of thousands of variants associated with complex traits and most of them fall within the non-coding regions, they may not be the causal ones. The development of high-throughput functional assays leads to the discovery of experimental validated non-coding functional variants. However, these validated variants are rare due to technical difficulty and financial cost. The small sample size of validated variants makes it less reliable to develop a supervised machine learning model for achieving a whole genome-wide prediction of non-coding causal variants.Results: We will exploit a deep transfer learning model, which is based on convolutional neural network, to improve the prediction for functional non-coding variants (NCVs). To address the challenge of small sample size, the transfer learning model leverages both large-scale generic functional NCVs to improve the learning of low-level features and context-specific functional NCVs to learn high-level features toward the context-specific prediction task. By evaluating the deep transfer learning model on three MPRA datasets and 16 GWAS datasets, we demonstrate that the proposed model outperforms deep learning models without pretraining or retraining. In addition, the deep transfer learning model outperforms 18 existing computational methods in both MPRA and GWAS datasets.