Sparse Convolutional Denoising Autoencoders for Genotype Imputation

Sparse Convolutional Denoising Autoencoders for Genotype Imputation
复制标题

DOI:
10.3390/genes10090652
复制
发表时间:
2019-09-01
期刊:
影响因子:
3.5
通讯作者:
Shi, Xinghua
Shi, Xinghua
中科院分区:
生物学3区
文献类型:
--
作者:
Chen, Junjie;Shi, Xinghua

文献摘要

被引文献

相似文献

基因型插补(可以通过计算插补缺失的基因型)是基因组分析中的重要工具,范围从全基因组关联到表型预测。传统的基因型插补方法通常基于单倍型聚类算法、隐马尔可夫模型 (HMM) 和统计推断。最近有报道称,基于深度学习的方法可以适当解决各个领域的数据丢失问题。为了探索深度学习在基因型插补方面的性能,在本研究中,我们提出了一种称为稀疏卷积去噪自动编码器(SCDA)的深度模型来插补缺失的基因型。我们使用卷积层构建了 SCDA 模型,该卷积层可以提取基因型数据中的各种相关性或连锁模式,并应用 L-1 正则化产生的稀疏权重矩阵来处理高维数据。我们综合评估了 SCDA 模型在不同场景下的性能,分别对酵母和人类基因型数据进行基因型插补。我们的结果表明 SCDA 具有很强的鲁棒性,并且显着优于流行的无参考插补方法。因此,这项研究指出了深度学习模型在基因组研究中缺失数据插补的另一种新颖应用。
Genotype imputation, where missing genotypes can be computationally imputed, is an essential tool in genomic analysis ranging from genome wide associations to phenotype prediction. Traditional genotype imputation methods are typically based on haplotype-clustering algorithms, hidden Markov models (HMMs), and statistical inference. Deep learning-based methods have been recently reported to suitably address the missing data problems in various fields. To explore the performance of deep learning for genotype imputation, in this study, we propose a deep model called a sparse convolutional denoising autoencoder (SCDA) to impute missing genotypes. We constructed the SCDA model using a convolutional layer that can extract various correlation or linkage patterns in the genotype data and applying a sparse weight matrix resulted from the L-1 regularization to handle high dimensional data. We comprehensively evaluated the performance of the SCDA model in different scenarios for genotype imputation on the yeast and human genotype data, respectively. Our results showed that SCDA has strong robustness and significantly outperforms popular reference-free imputation methods. This study thus points to another novel application of deep learning models for missing data imputation in genomic studies.