Imputation of single-cell gene expression with an autoencoder neural network

Imputation of single-cell gene expression with an autoencoder neural network
复制标题

DOI:
10.1007/s40484-019-0192-7
复制
发表时间:
2020-03-01
影响因子:
3.1
通讯作者:
Fu, Audrey Qiuyan
Fu, Audrey Qiuyan
中科院分区:
生物学4区
文献类型:
--
作者:
Badsha, Md Bahadur;Li, Rui;Fu, Audrey Qiuyan

文献摘要

被引文献

相似文献

背景单细胞RNA测序(scRNA-seq)是一种快速发展的技术,能够以前所未有的分辨率测量基因表达水平。尽管可以通过单个实验检测的细胞数量呈爆炸性增长,但scRNA-seq仍然有几个局限性,包括高脱落率,这导致大量基因在scRNA-seq数据中具有零读段计数,并使下游分析复杂化。具体来说,我们的后期(使用AutoEncoder学习)方法使用参数的随机初始值训练自动编码器,而我们的TRANSLATE结果在模拟和真实的数据上,LATE和TRANSLATE优于现有的scRNA-seq插补方法,在大多数情况下实现较低的均方误差,恢复非线性基因-基因关系,并更好地分离细胞类型。他们也是高度可扩展的,可以有效地处理超过100万个细胞在短短几个小时的GPU.ConclusionsWe证明,我们的非参数方法的基础上自动编码器的插补是强大的,高效的。
BackgroundSingle-cell RNA-sequencing (scRNA-seq) is a rapidly evolving technology that enables measurement of gene expression levels at an unprecedented resolution. Despite the explosive growth in the number of cells that can be assayed by a single experiment, scRNA-seq still has several limitations, including high rates of dropouts, which result in a large number of genes having zero read count in the scRNA-seq data, and complicate downstream analyses.MethodsTo overcome this problem, we treat zeros as missing values and develop nonparametric deep learning methods for imputation. Specifically, our LATE (Learning with AuToEncoder) method trains an autoencoder with random initial values of the parameters, whereas our TRANSLATE (TRANSfer learning with LATE) method further allows for the use of a reference gene expression data set to provide LATE with an initial set of parameter estimates.ResultsOn both simulated and real data, LATE and TRANSLATE outperform existing scRNA-seq imputation methods, achieving lower mean squared error in most cases, recovering nonlinear gene-gene relationships, and better separating cell types. They are also highly scalable and can efficiently process over 1 million cells in just a few hours on a GPU.ConclusionsWe demonstrate that our nonparametric approach to imputation based on autoencoders is powerful and highly efficient.